Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models

T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper claims that the Poisson Hierarchical Indian Buffet Process reduces the joint marginal distribution of sparse grouped counts to an exact compound Poisson representation, enabling tractable posterior and predictive inference for…

desk verdict Sound and original theory for hierarchical Poisson IBPs, but the advertised structural-zero interpretation is not realized by the fitted gamma/GG models, and the empirical validation is too thin to support the microbiome claims. read the letter →

arxiv 2502.01919 v2 pith:ILERDUVJ submitted 2025-02-04 stat.ML cs.LGmath.PRmath.STstat.TH

classification stat.MLcs.LGmath.PRmath.STstat.TH MSC 60G0962F1560G5762P1060C05
keywords BayesiannonparametricsHierarchicalIndianBuffetProcessspeciessamplingmodelsmicrobiomecountdatazero-inflationunseencompoundPoissongeneralizedgamma
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces the Poisson Hierarchical Indian Buffet Process (PHIBP), a Bayesian nonparametric prior for sparse, grouped count data in which an infinite catalogue of species is shared across groups through a hierarchy of Poisson random measures. It establishes that the joint marginal distribution of the group-sum counts is exactly a compound Poisson process driven by a latent allocation process, giving a tractable exact sampling scheme and explicit posterior and prediction rules. The authors argue this provides a principled treatment of zeros in microbiome count matrices, distinguishing species that are truly absent from species that are present but undetected, and yields model-based diversity and unseen-species estimates. The practical claim is that a generalized gamma specification captures rare species better than a gamma prior, avoiding the artifacts of pseudocount-based methods.

What carries the argument

The species allocation process $A_J = (\sum_{l=1}^\infty \xi_{j,l}\,\delta_{Y_l}: j\in[J])$, where $\xi_{j,l}$ are Poisson counts of latent OTUs for species $Y_l$ in group $j$, is the hidden combinatorial engine of the paper. It converts the three-level hierarchy of Poisson random measures into a multivariate Poisson Indian buffet process, whose thinning leads to compound Poisson representations using zero-truncated Poisson (tP) and mixed truncated Poisson (MtP) variables, with Multinomial allocation of counts across groups. Finite Gibbs exchangeable partition functions $\Xi^{[n_{j,l}]}_{x_{j,l}}$ carry the posterior normalization, and the generalized gamma versus gamma choice of Lévy densities controls whether rare species are preserved in the frequency-of-frequency behavior.

What would settle it

Simulate microbiome-like data with a genuine extra zero-inflation layer, e.g., each observed count independently set to zero with probability p beyond the Poisson rate, then fit both the PHIBP and a PHIBP extended with such a zero-inflation component; if the extended model recovers p > 0 with substantially better predictive likelihood, the PHIBP's zero decomposition is not identified.

Watch

Extended reading notes

Core claim

The paper's central discovery is that the PHIBP marginal distribution, posterior, and predictive rules reduce to compound Poisson representations built from components already in the literature. Theorem 3.1 shows the sum process is distributionally equal to a compound Poisson process whose atoms are species with MtP-distributed latent OTU counts, with group allocation following a Multinomial distribution, enabling exact generative sampling of the marginal. Theorem 4.1 gives the posterior as a decomposition in which observed species have posterior rates split into contributions from OTUs not yet seen in the sample and contributions from observed OTUs, while unseen species retain a residual completely random measure. Proposition 4.4 decomposes prediction for a new sample into arrivals of completely new species, new OTUs of existing species, and additional counts for known OTUs, which together answer classical unseen-species questions in a form suited to sequencing data where the total number of reads is itself random.

Load-bearing premise

The model assumes that every zero in the observed count matrix is fully explained by the latent Poisson sampling rate, so there is no separate zero-inflation mechanism (such as PCR dropout) acting on counts; if such a mechanism exists, the split between structural and sampling zeros is not identified.

Editorial extensions

If this is right

  • Theorem 3.1 gives an exact generative sampler for the PHIBP marginal distribution, so the complex multi-dimensional count distributions can be simulated directly without approximating the infinite hierarchy.
  • The posterior representation splits each observed species' latent rate into an unobserved-OTU part and an observed-OTU part, which is what lets the model express uncertainty about zeros rather than collapsing them to zero.
  • Proposition 4.4 divides prediction into completely new species, new OTUs of known species, and additional counts for known OTUs, yielding answers to unseen-species questions under random total reads.
  • On the real microbiome dataset, the generalized gamma PHIBP matches the power-law frequency-of-frequency distribution of the test data and gives higher predictive likelihood than the gamma PHIBP, which underestimates rare species and inflates beta diversity.
  • The same framework extends to covariates, technical variation adjustments, and strain-trait modeling, with the latent OTU components providing interpretable fitness rates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the construction uses only Poisson random measures and zero-truncated Poisson variables, the same compound-Poisson machinery should transfer to other sparse count domains such as genetic variant discovery or text token frequencies without new theory.
  • When exact sequence variants are observed, the latent OTU counts can be treated as fixed inputs, turning the model into a direct hierarchical sampler for strain-level fitness rates; this is an extension the paper outlines but does not develop empirically.
  • The predictive unseen entropy $U_{j,M_{j+1}}$ could serve as a sampling-design criterion, indicating which group is expected to yield the most diversity among species not yet seen in any sample.
  • One could test the zero-handling claim by fitting the PHIBP alongside a version with an explicit extra zero-inflation layer on synthetic data generated with known dropout; the paper does not include such a comparison.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces the Poisson Hierarchical Indian Buffet Process (PHIBP), a three-level CRM construction for grouped sparse count data: B0 ~ CRM(τ0,F0), Bj | B0 ~ CRM(τj,B0), and observations given as Poisson random measures with intensities γ_{i,j}B_j. The central results are a compound-Poisson representation and exact sampling scheme for the marginal sum process (Theorem 3.1), posterior characterizations (Theorem 4.1 and Propositions 4.1–4.3), and a three-component predictive rule for new samples (Proposition 4.4). The authors propose Bayesian diversity measures and an unseen-species entropy (Sections 2.2 and 4.6), and they compare generalized-gamma and gamma versions of the model on simulated data and on a 12-sample microbiome count dataset (Section 6). The main theoretical derivations appear internally consistent and are developed through standard Poisson-process and subordinator calculus, with proofs in Appendix A; the main substantive problem is the claimed distinction between structural and sampling zeros, which the fitted model does not actually implement.

Significance. If the mathematical claims are correct, the paper provides a useful unifying treatment of hierarchical Poisson IBP models, with explicit joint distributions, exact marginal sampling, and posterior samplers, extending earlier work of James and others in the species-sampling literature. The compound-Poisson representation and the Gibbs-partition interpretation in Proposition 4.1 are genuine technical contributions, and the GG-specific formulas connecting to Stirling numbers are useful. The empirical part is suggestive but limited: it compares only two versions of the same model, and the advertised practical contribution of distinguishing technical from biological zeros is not supported by the model's posterior structure. The paper's potential significance therefore rests on the theoretical development rather than on the microbiome-specific zero-inflation claims, which currently need correction.

major comments (3)
  1. [Section 2.1 and Proposition 4.2, Eq. (4.5)] The advertised distinction between structural zeros and sampling zeros is not realized by the fitted gamma/GG PHIBP. For any species observed in at least one group, H_l > 0 almost surely, and Proposition 4.2 gives the posterior local rate as σ̃_{j,l}(H_l) = σ̂_{j,l}(H_l) + Σ_{k=1}^{X_{j,l}} S_{j,k,l}. For the gamma and generalized gamma Levy densities used in Section 6, σ̂_{j,l}(H_l) is the value of an infinite-activity subordinator after an Esscher transform; in the gamma case it is Gamma(θ_j H_l, ζ_j + Σ_i γ_{i,j}), which has no atom at zero. Hence Pr(σ̃_{j,l}(H_l)=0 | N_{j,l}=0) = 0. A zero count for a globally observed species is always a sampling zero under the implemented model, despite Section 2.1 claiming that the posterior 'reflects whether a zero count Nj,l = 0 likely corresponds to a structural zero (where the posterior for the latent rate σj,l becomes concentrated at zero)'. The posterior can be small, but it cannot represent biological absence, so the structural-versus-sampling-zero decomposition and the diversity/unseen-species interpretations built on it in Sections 2.2 and 4.6 are not supported by the model actually fitted. The authors should either add an explicit zero-inflation layer that places prior mass at zero or substantially reframe these claims.
  2. [Section 6.2 and Appendix B] The empirical evaluation does not test the zero-handling claim. The experiments compare GG PHIBP against Gamma PHIBP, but neither prior places any mass at σ_{j,l}=0, so differences in frequency-of-frequency distributions, test log-likelihoods, and diversity measures cannot validate the structural-zero mechanism advertised in the abstract and Section 2.1. A comparison against a model with an explicit zero-inflation component, or a direct posterior check for concentration near zero versus positive mass, would be needed to support the claim that the framework 'explicitly distinguish[es] between technical and biological zeros'. As it stands, the microbiome experiments only show that the generalized-gamma subordinator fits the power-law frequency-of-frequency pattern better than the gamma subordinator, which is a weaker and more conventional claim.
  3. [Section 4.6 and Proposition 4.5] The unseen-species entropy U_{j,M_j+1} is presented as a novel contribution, but its identification rests on the same problematic zero decomposition. The quantity is defined on the posterior rates of completely new species, and Proposition 4.5 provides a sampling algorithm for its scaled version. However, the accompanying discussion in Section 4.6 connects this quantity to the number of previously unobserved species Q2 and to the paper's broader claims about handling sparse co-occurrence. Because the model has no atom at zero, the 'unseen' component can only describe low-rate species, not species that are truly absent from the community. The text should either clarify that the model addresses rarity rather than absence or add a mechanism that can represent genuine structural zeros.
minor comments (4)
  1. [Appendix A and Remark A.3] The proof of Theorem 4.1 refers to 'Proposition 5.2' where Proposition 4.2 is meant, and the same cross-reference error appears in Remark A.3.
  2. [Section 3.4, Eq. (3.7)] Equation (3.7) is mis-rendered: as printed, the displayed expression is not a well-formed joint density, and the notation ψ^{(c_{j,k,l})}_j is introduced only later in the text. The formula should be rewritten or the notation defined immediately.
  3. [Appendix A, proof of Theorem 3.1] The sentence 'We next treat the case of Theorem 3.3' appears to refer to Theorem 3.1, and there are several typos such as 'Propostion' and 'Thoerem' in the appendix; these should be corrected before publication.
  4. [Section 1 and Section 6] The text states that the framework is accompanied by 'freely available resources to validate our model's performance', but no repository or URL is provided in the manuscript; a link or reference to public code would improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Theorem 3.1 is a derived compound-Poisson representation from CRM thinning identities, and the empirical claims are held-out posterior predictive checks.

full rationale

The central derivation chain is self-contained conditional on standard CRM/Poisson-IBP identities. Theorem 3.1 is not assumed: it is obtained in Appendix A by applying [16]'s zero-set thinning decomposition to the random base measure B0, yielding the Allocation process AJ and then the compound-Poisson representation (3.9); the proof uses only infinite divisibility of subordinators and Poisson thinning, not the theorem's conclusion. Proposition 4.4 and Theorem 4.1 are consequences of the same decomposition, not fitted quantities. The empirical section compares GG vs Gamma PHIBP on held-out test splits using predictive likelihoods and FoF statistics, so no fitted constant is renamed as a prediction. The main self-citations ([16], [19]) point to peer-reviewed, parameter-free mathematical identities with stated assumptions, and [18] is explicitly described as preliminary and non-load-bearing ('all results derived in a fully self-contained and updated manner'). The structural-zero/sampling-zero distinction flagged in Section 2.1 is a model-identification/correctness concern about the absence of an atom at zero in the gamma/GG posterior, not a circular dependency: the model's zero decomposition does not feed back into the derivation of Theorem 3.1. Accordingly, no circular step is exhibited.

Assumptions & free parameters 4 free parameters · 5 assumptions · 2 invented entities

The central theoretical claims rest on standard Poisson calculus and CRM theory from the literature. The model-specific assumptions are the mixed-Poisson hierarchy and the chosen Levy measures. The empirical application introduces free parameters inferred from data and fixes others by hand. No new physical entities are postulated; the invented entities are latent stochastic processes internal to the model.

free parameters (4)
  • α0, θ0 = inferred via MCMC
    Generalized gamma base-measure parameters. In experiments they are estimated from data; the theoretical results do not depend on specific values.
  • αj, θj for j∈[J] = inferred via MCMC
    Group-level generalized gamma Levy parameters estimated from training data in experiments.
  • ζ0, ζj = fixed to 1
    Chosen by hand for simplicity in experiments; not inferred, despite being mentioned as inferable.
  • γi,j = fixed to 1
    Sample-specific scaling factors fixed to 1 in all experiments, ignoring sequencing depth differences across the 12 microbiome samples.
assumptions (5)
  • standard math Poisson random measure and CRM theory: joint jumps are Poisson points with mean measure (A.5), from Sato (2013) Theorem 30.1.
    Used in Appendix A.2 to establish the multivariate CRM representation underlying the posterior decomposition.
  • standard math Multivariate IBP results of James (2017) and Pitman (1997) for zero-truncated Poisson and compound Poisson representations.
    Repeatedly invoked for Proposition 3.1, Theorem 3.1, and Proposition 4.3.
  • standard math Finite Gibbs EPPF results from Pitman (2006), Kolchin (1986), and James, Lijoi, Prünster (2009).
    Used to identify conditional distributions in Proposition 4.1 and posterior rate representations.
  • domain assumption Mixed-Poisson model assumption: observed counts are conditionally independent Poisson given latent CRM intensities.
    This is the defining model, not derived. The microbiome interpretation of zeros depends on this assumption.
  • ad hoc to paper Generalized gamma Levy density τ(s) = θ/(Γ(1-α)) s^{-α-1} e^{-ζs} is the relevant species abundance model.
    Chosen for power-law behavior; the paper's GG versus Gamma comparison is conditioned on this specific choice.
invented entities (2)
  • Allocation process A_J
    purpose: A multivariate Poisson IBP tracking species and OTU multiplicities across groups; central organizing object for exact sampling and prediction.
    A latent stochastic process introduced to decompose the PHIBP. It has no external falsifiable handle outside the model itself.
  • Latent OTU frequency measure F_{j,l}
    purpose: Represents OTU-level relative abundances within a species, used to interpret technical versus biological zeros.
    An interpretive construct from the model's latent variables; cannot be independently verified without OTU-resolution data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models." pith.science (2026). https://pith.science/paper/ILERDUVJ

@misc{pith2026250201919,
  author       = {Pith},
  title        = {Pith review of: Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ILERDUVJ}},
  note         = {Machine review of arXiv:2502.01919}
}
read the original abstract

We introduce the Poisson Hierarchical Indian Buffet Process (PHIBP), a new class of species sampling models designed to address the challenges of complex, sparse count data by facilitating information sharing across and within groups. Our theoretical developments enable a tractable Bayesian nonparametric framework with machine learning elements, accommodating a potentially infinite number of species (taxa) whose parameters are learned from data. Focusing on microbiome analysis, we address key gaps by providing a flexible multivariate count model that accounts for overdispersion and robustly handles diverse data types (OTUs, ASVs). We introduce novel parameters reflecting species abundance and diversity. The model borrows strength across groups while explicitly distinguishing between technical and biological zeros to interpret sparse co-occurrence patterns. This results in a framework with tractable posterior inference, exact generative sampling, and a principled solution to the unseen species problem. We describe extensions where domain experts can incorporate knowledge through covariates and structured priors, with potential for strain-level analysis. While motivated by ecology, our work provides a broadly applicable methodology for hierarchical count modeling in genetics, commerce, and text analysis, and has significant implications for the broader theory of species sampling models arising in probability and statistics.

Figures

Figures reproduced from arXiv: 2502.01919 by the authors.

Figure 1
Figure 1. Simulated data experiments. Left and middle: Posterior samples for [PITH_FULL_IMAGE:figures/full_fig_p021_1.png] view at source ↗
Figure 2
Figure 2. Left and Middle: Comparison of FoF distributions of the simulated data (under [PITH_FULL_IMAGE:figures/full_fig_p023_2.png] view at source ↗
Figure 3
Figure 3. Left and Middle: posterior distributions of [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Posterior distributions of the GG HIBP model parameters inferred from the [PITH_FULL_IMAGE:figures/full_fig_p033_4.png]
Figure 5
Figure 5. Figure 5: Posterior distributions of the Gamma HIBP model parameters inferred from the [PITH_FULL_IMAGE:figures/full_fig_p033_5.png]
Figure 6
Figure 6. Figure 6: Test log-likelihood values of GG and Gamma HIBP models for microbiome data. [PITH_FULL_IMAGE:figures/full_fig_p034_6.png]
Figure 7
Figure 7. Figure 7: Posterior distributions of the GG HIBP model parameters inferred from the [PITH_FULL_IMAGE:figures/full_fig_p034_7.png]
Figure 8
Figure 8. Figure 8: Posterior distributions of the Gamma HIBP model parameters inferred from the [PITH_FULL_IMAGE:figures/full_fig_p035_8.png]
Figure 9
Figure 9. Figure 9: KS statistics between the FoFs of the predicted data and the observed test data. [PITH_FULL_IMAGE:figures/full_fig_p035_9.png]
Figure 10
Figure 10. Figure 10: Posterior α diversities measured as Shannon entropies for the microbiome data. [9] Fisher, R. A., Corbet, S. A. and Williams, C. B. [1943], ‘The relation between the number of species and the number of individuals in a random sample of an animal population.’, The Jour…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Limit Theorems for the Pitman-Yor Frequency Spectrum

    math.PR 2026-07 accept novelty 6.5 of 10

    For Pitman–Yor partitions, sums ∑_{j=⌊λn⌋}^{⌊μn⌋} M_jn converge in law (conditionally and marginally) to explicit mixture distributions built from truncated stable subordinators.

Reference graph

Works this paper leans on

52 extracted references · 49 canonical work pages · cited by 1 Pith paper

  1. [1]

    and Naulet, Z

    Balocchi, C., Favaro, S. and Naulet, Z. [2024], ‘Bayesian nonparametric inference for” species- sampling” problems’, Statistical Science, to appear . To Appear

  2. [2]

    E., Pedersen, J

    Barndorff-Nielsen, O. E., Pedersen, J. and Sato, K.-i. [2001], ‘Multivariate subordination, self- decomposability and stability’, Advances in Applied Probability 33(1), 160–187

  3. [3]

    and Favaro, S

    Beraha, M. and Favaro, S. [2023], ‘Transform-scaled process priors for trait allocations in Bayesian nonparametrics’, arXiv preprint arXiv:2303.17844

  4. [4]

    and Jordan, M

    Blei, D., Ng, A. and Jordan, M. [2003], ‘Latent Dirichlet Allocation. ’, Journal of Machine Learn- ing Research 3, 993–1022

  5. [5]

    Brooks, S. P. and Gelman, A. [1997], ‘General methods for monitoring convergence of iterative simulations’, Journal of Computational and Graphical Statistics 7, 434–455. 34 Fig 6: Test log-likelihood values of GG and Gamma HIBP models for microbiome data. Fig 7: Posterior distributions of the GG HIBP model parameters inferred from the microbiome data, alo...

  6. [6]

    [2016], ‘The Ubiquitous Ewens Sampling Formula’, Statistical Science 31(1), 1 – 19

    Crane, H. [2016], ‘The Ubiquitous Ewens Sampling Formula’, Statistical Science 31(1), 1 – 19

  7. [7]

    Favaro, S., Lijoi, A., Mena, R. H. and Pr¨ unster, I. [2009], ‘Bayesian non-parametric inference for species variety with a two-parameter poisson–dirichlet process prior’, JRSSB 71(5), 993– 1008

  8. [8]

    Ferguson, T. S. [1973], ‘A Bayesian Analysis of Some Nonparametric Problems’, Ann. Statist. 1, 209–230. 36 Group 1 GG 5.23+-0.024 GA 5.159+-0.021 Group 2 GG 5.103+-0.028 GA 5.01+-0.026 Group 3 GG 6.219+-0.019 GA 6.043+-0.016 Group 4 GG 6.108+-0.023 GA 5.913+-0.017 Group 5 GG 6.2+-0.019 GA 6.033+-0.015 Group 6 GG 6.192+-0.017 GA 6.049+-0.014 Group 7 GG 6.1...

Show all 52 references
  1. [9]

    A., Corbet, S

    Fisher, R. A., Corbet, S. A. and Williams, C. B. [1943], ‘The relation between the number of species and the number of individuals in a random sample of an animal population.’, The Journal of Animal Ecology pp. 42–58

  2. [10]

    and De Iorio, M

    Franzolini, B., Cremaschi, A., van den Boom, W. and De Iorio, M. [2023], ‘Bayesian cluster- ing of multiple zero-inflated outcomes’, Philosophical Transactions of the Royal Society A 381(2247), 20220145

  3. [11]

    and Rubin, B

    Gelman, A. and Rubin, B. [1992], ‘Inference from iterative simulation using multiple sequences’, Statistical Science 7, 457–511

  4. [12]

    and Griffiths, T

    Ghahramani, Z. and Griffiths, T. [2005], ‘Infinite latent feature models and the Indian buffet process’, Advances in neural information processing systems 18

  5. [13]

    and James, L

    Ishwaran, H. and James, L. F. [2001], ‘Gibbs sampling methods for stick-breaking priors.’, J. Amer. Stat. Assoc. 96, 161–173

  6. [14]

    and James, L

    Ishwaran, H. and James, L. F. [2003], ‘Generalized weighted Chinese restaurant processes for species sampling mixture models.’, Statist. Sinica 13, 1211–1235

  7. [15]

    James, L. F. [2002], Poisson process partition calculus with applications to exchangeable mod- els and Bayesian nonparametrics, Technical report, Unpublished manuscript. ArXiv math. PR/0205093

  8. [16]

    James, L. F. [2017], ‘Bayesian Poisson Calculus for Latent Feature Modeling via Generalized Indian Buffet Process Priors’, Ann. Stat. 45, 2016–2045

  9. [17]

    James, L. F. [2019], ‘Stick-breaking Pitman-Yor processes given the species sampling size’, arXiv preprint arXiv:1908.07186

  10. [18]

    F., Lee, J

    James, L. F., Lee, J. and Pandey, A. [2023], ‘Bayesian analysis of generalized hierarchical indian buffet processes for within and across group sharing of latent features’, arXiv preprint arXiv:2304.05244

  11. [19]

    F., Lijoi, A

    James, L. F., Lijoi, A. and Pr¨ unster, I. [2009], ‘Posterior analysis for normalized random measures with independent increments.’, Scand. J. Stat. 36, 76–97

  12. [20]

    and Holmes, S

    Jeganathan, P. and Holmes, S. P. [2021], ‘A statistical perspective on the challenges in molec- ular microbial biology’, Journal of Agricultural, Biological and Environmental Statistics 26(2), 131–160. PHIBP 37

  13. [21]

    Jones, P. C. T., Mollison, J. E. and Quenouille, M. H. [1948], ‘A technique for the quantitative estimation of soil micro-organisms’, Journal of General Microbiology 2(1), 54–69

  14. [22]

    Kingman, J. F. [1975], ‘Random discrete distributions’, JRSSB 37(1), 1–15

  15. [23]

    Kolchin, V. F. [1986], Random mappings, Translation Series in Mathematics and Engineering, Optimization Software Inc. Publications Division, New York

  16. [24]

    Lee, M. D. [n.d.], ‘A full example workflow for amplicon data’, https://astrobiomike.github.io/ amplicon/dada2 workflow ex

  17. [25]

    D., Walworth, N

    Lee, M. D., Walworth, N. G., Sylvan, J. B., Edwards, K. J. and Orcutt, B. N. [2015], ‘Microbial communities on seafloor basalts at dorado outcrop reflect level of alteration and highlight global lithic clades’, Frontiers in Microbiology 6, 1470

  18. [26]

    Lijoi, A., Mena, R. H. and Pr¨ unster, I. [2007], ‘Bayesian nonparametric estimation of the prob- ability of discovering new species’, Biometrika 94(4), 769–786

  19. [27]

    Lo, A. Y. [1982], ‘Bayesian nonparametric statistical inference for Poisson point processes’, Zeitschrift f¨ ur Wahrscheinlichkeitstheorie und verwandte Gebiete59(1), 55–66

  20. [28]

    and Broderick, T

    Masoero, L., Camerlenghi, F., Favaro, S. and Broderick, T. [2018], ‘Posterior representations of hierarchical completely random measures in trait allocation models.’, BNP-NeurIPS. 2018

  21. [29]

    and Broderick, T

    Masoero, L., Camerlenghi, F., Favaro, S. and Broderick, T. [2022], ‘More for less: predict- ing and maximizing genomic variant discovery via Bayesian nonparametrics’, Biometrika 109(1), 17–32

  22. [30]

    and Møller, J

    McCullagh, P. and Møller, J. [2006], ‘The permanental process’, Advances in Applied Probability 38(4), 873–888

  23. [31]

    [1996], ‘Some developments of the Blackwell-MacQueen urn scheme’, Statistics, prob- ability and game theory IMS Lecture Notes Monogr

    Pitman, J. [1996], ‘Some developments of the Blackwell-MacQueen urn scheme’, Statistics, prob- ability and game theory IMS Lecture Notes Monogr. Ser., 30, Inst. Math. Statist., Hayward, CA, 245–267

  24. [32]

    [1997], ‘Partition structures derived from Brownian motion and stable subordinators’, Bernoulli 3, 79–96

    Pitman, J. [1997], ‘Partition structures derived from Brownian motion and stable subordinators’, Bernoulli 3, 79–96

  25. [33]

    [2003], Poisson-Kingman partitions, in R

    Pitman, J. [2003], Poisson-Kingman partitions, in R. Goldstein, ed., ‘Institute of Mathematical Statistics’, Hayward, California, pp. 1–34

  26. [34]

    [2006], Combinatorial stochastic processes, Lectures from the 32nd Summer School on Probability Theory held in Saint-Flour, July 7–24, 2002

    Pitman, J. [2006], Combinatorial stochastic processes, Lectures from the 32nd Summer School on Probability Theory held in Saint-Flour, July 7–24, 2002. With a foreword by Jean Picard. Lecture Notes in Mathematics, 1875, Springer-Verlag, Berlin

  27. [35]

    [2017], Mixed Poisson and negative binomial models for clustering and species sampling- Unpublished manuscript

    Pitman, J. [2017], Mixed Poisson and negative binomial models for clustering and species sampling- Unpublished manuscript

  28. [36]

    and Yor, M

    Pitman, J. and Yor, M. [1997], ‘The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator’, Ann. Probab. 25, 855–900

  29. [37]

    K., Stephens, M

    Pritchard, J. K., Stephens, M. and Donnelly, P. [2000], ‘Inference of population structure using multilocus genotype data’, Genetics 155(2), 945–959

  30. [38]

    and Trippa, L

    Ren, B., Bacallado, S., Favaro, S., Holmes, S. and Trippa, L. [2017], ‘Bayesian nonparametric or- dination for the analysis of microbial communities’,J. Amer. Statist. Assoc. 112(520), 1430– 1442

  31. [39]

    and Pavoine, S

    Ricotta, C., Szeidl, L. and Pavoine, S. [2021], ‘Towards a unifying framework for diversity and dissimilarity coefficients’, Ecological Indicators 129, 107971

  32. [40]

    and Holmes, S

    Sankaran, K. and Holmes, S. P. [2019], ‘Latent variable modeling for the microbiome’, Biostatis- tics 20(4), 599–614

  33. [41]

    Sankaran, K., Kodikara, S., Li, J. J. and Le Cao, K.-A. [2024], ‘Semisynthetic simulation for microbiome data analysis’, bioRxiv pp. 2024–10

  34. [42]

    [2013], L´ evy Processes and Infinitely Divisible Distributions, number 68 in ‘Cambridge Studies in Advanced Mathematics’, 2nd edn, Cambridge University Press, Cambridge

    Sato, K.-i. [2013], L´ evy Processes and Infinitely Divisible Distributions, number 68 in ‘Cambridge Studies in Advanced Mathematics’, 2nd edn, Cambridge University Press, Cambridge

  35. [43]

    [2019], Allocative Poisson Factorization for Computational Social Science, PhD thesis, University of Massachusetts Amherst, 2019

    Schein, A. [2019], Allocative Poisson Factorization for Computational Social Science, PhD thesis, University of Massachusetts Amherst, 2019

  36. [44]

    [2021], ‘The magical Ewens sampling formula’, Bulletin of the London Mathematical Society 53(6), 1563–1582

    Tavar´ e, S. [2021], ‘The magical Ewens sampling formula’, Bulletin of the London Mathematical Society 53(6), 1563–1582

  37. [45]

    W., Jordan, M

    Teh, Y. W., Jordan, M. I., Beal, M. J. and Blei, D. M. [2006], ‘Hierarchical Dirichlet Processes.’, J. Amer. Statist. Assoc. 101, 1566–1581

  38. [46]

    and Jordan, M

    Thibaux, R. and Jordan, M. I. [2007], Hierarchical beta processes and the Indian buffet process, in ‘International conference on artificial intelligence and statistics’, pp. 564–571

  39. [47]

    Titsias, M. K. [2008], ‘The infinite gamma-Poisson feature model’, Advances in Neural Informa- tion Processing Systems. 2008. 38

  40. [48]

    and Caron, F

    Todeschini, A., Miscouridou, X. and Caron, F. [2020], ‘Exchangeable random measures for sparse and modular graphs with overlapping communities’,Journal of the Royal Statistical Society: Series B (Statistical Methodology) 82(2), 487–520

  41. [49]

    Willis, A. D. and Martin, B. D. [2022], ‘Estimating diversity in networked ecological communi- ties’, Biostatistics 23(1), 207–222

  42. [50]

    and Carin, L

    Zhou, M. and Carin, L. [2015], ‘Negative Binomial Process Count and Mixture Modeling’, IEEE Transactions on Pattern Analysis and Machine Intelligence 37, 307–320

  43. [51]

    and Walker, S

    Zhou, M., Favaro, S. and Walker, S. G. [2017], ‘Frequency of frequencies distributions and size- dependent exchangeable random partitions’, J. Amer. Statist. Assoc. 112(520), 1623–1635

  44. [52]

    Zhou, M., Madrid-Padilla, O. H. and Scott, J. G. [2016], ‘Priors for Random Count Matrices Derived from a Family of Negative Binomial Processes.’,J. Amer. Statist. Assoc. 111, 1144– 1156

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.