Pith. sign in

REVIEW 1 cited by

Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.01919 v2 pith:ILERDUVJ submitted 2025-02-04 stat.ML cs.LGmath.PRmath.STstat.TH

Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models

classification stat.ML cs.LGmath.PRmath.STstat.TH
keywords speciessamplinganalysiscountdatahierarchicalmodelsacross
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce the Poisson Hierarchical Indian Buffet Process (PHIBP), a new class of species sampling models designed to address the challenges of complex, sparse count data by facilitating information sharing across and within groups. Our theoretical developments enable a tractable Bayesian nonparametric framework with machine learning elements, accommodating a potentially infinite number of species (taxa) whose parameters are learned from data. Focusing on microbiome analysis, we address key gaps by providing a flexible multivariate count model that accounts for overdispersion and robustly handles diverse data types (OTUs, ASVs). We introduce novel parameters reflecting species abundance and diversity. The model borrows strength across groups while explicitly distinguishing between technical and biological zeros to interpret sparse co-occurrence patterns. This results in a framework with tractable posterior inference, exact generative sampling, and a principled solution to the unseen species problem. We describe extensions where domain experts can incorporate knowledge through covariates and structured priors, with potential for strain-level analysis. While motivated by ecology, our work provides a broadly applicable methodology for hierarchical count modeling in genetics, commerce, and text analysis, and has significant implications for the broader theory of species sampling models arising in probability and statistics.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Limit Theorems for the Pitman-Yor Frequency Spectrum

    math.PR 2026-07 accept novelty 6.5

    For Pitman–Yor partitions, sums ∑_{j=⌊λn⌋}^{⌊μn⌋} M_jn converge in law (conditionally and marginally) to explicit mixture distributions built from truncated stable subordinators.