{"id":"d2686326-86d5-4e4e-af97-6a2170e722ec","arxiv_id":"2502.01919","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new hierarchical Bayesian nonparametric model for sparse count data provides exact generative sampling, tractable posterior inference, and predictive rules for unseen species in microbiome studies.","lead":"This paper introduces the Poisson Hierarchical Indian Buffet Process, a Bayesian model for sparse, grouped count data where an infinite catalog of species is shared across groups. It provides exact sampling and prediction formulas, then applies the model to microbiome counts to estimate diversity and unseen species.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 2.1's structural-vs-sampling-zero distinction is not realized by the gamma/GG PHIBP: the posterior local rate has no atom at zero, so biological absence is unrepresentable for observed species.","rationale":"The reader's conditional verdict is well placed: the main mathematical construction in Theorem 3.1 and the compound-Poisson sampling scheme appear correct, and the absence of external baselines, code, and an explicit zero-inflation comparison justifies conditional acceptance. The reader's weakest assumption concerned an external zero-inflation mechanism; the present stress test identifies a related but more internal issue. In the gamma and GG cases actually fitted in Section 6, the subordinators are infinite-activity, so every species that is observed anywhere has a positive local rate in every group almost surely. Consequently, the posterior for ̃σ_{j,l} given a zero count is a continuous distribution on (0,∞) with no point mass at zero. The paper's language in Section 2.1 of distinguishing structural zeros via a posterior 'concentrated at zero' is therefore an overstatement: the model can shrink rates toward zero but cannot certify absence. This does not invalidate Theorem 3.1, but it does weaken a headline applied contribution, supporting the conditional verdict and the need for either a revised interpretation or a model extension with an atom at zero. The proposed analytical check settles the issue definitively by computing the posterior distribution of the local rate at zero. Since the reader already recommended conditional acceptance on related grounds, the verdict should remain unchanged rather than move to accept or reject.","tokens_in":34826,"tokens_out":32564,"duration_ms":339617,"concrete_test":"Analytically evaluate the posterior of ̃σ_{j,l}(H_l) conditional on N_{j,l}=0 using Theorem 4.1 and Proposition 4.2 for the gamma case: it is Gamma(θ_j H_l, ζ_j + Σ_i γ_{i,j}), so Pr(̃σ_{j,l}=0 | N_{j,l}=0)=0; perform the corresponding check for the generalized gamma case. If both posteriors have no atom at zero, Section 2.1's structural-zero claim requires either an explicit atom-at-zero mechanism at the local-rate level or a revised interpretation in terms of shrinkage rather than exact absence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The algebraic core of Theorem 3.1 appears internally consistent, so the load-bearing problem lies in the manuscript's zero-decomposition claim. Section 2.1 states that a zero count can be inferred as a structural zero when the posterior of the latent rate σ_{j,l} is 'concentrated at zero', versus a sampling zero when the posterior keeps mass at positive values. But for every species observed anywhere in the data, H_l > 0 a.s., and the posterior local rate is (Proposition 4.2, Eq. 4.5) ̃σ_{j,l}(H_l) = ̂σ_{j,l}(H_l) + Σ_{k=1}^{X_{j,l}} S_{j,k,l}. For the gamma and generalized gamma subordinators used in Section 6, ̂σ_{j,l}(H_l) is an Esscher-transformed infinite-activity subordinator: e.g., in the gamma case it is Gamma(θ_j H_l, ζ_j + Σ_i γ_{i,j}), whose distribution is absolutely continuous on (0,∞) and has zero probability at 0. Hence Pr(̃σ_{j,l}(H_l)=0 | N_{j,l}=0) = 0: a zero count for a globally observed species is always a sampling zero under the implemented model. The posterior can shrink the rate toward zero, but it cannot represent a true biological absence, so the advertised distinction between structural and sampling zeros, and the diversity/unseen-species interpretations built on it in Sections 2.2 and 4.6, are not supported by the fitted gamma/GG model.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Poisson Hierarchical Indian Buffet Process (PHIBP), a three-level CRM construction for grouped sparse count data: B0 ~ CRM(τ0,F0), Bj | B0 ~ CRM(τj,B0), and observations given as Poisson random measures with intensities γ_{i,j}B_j. The central results are a compound-Poisson representation and exact sampling scheme for the marginal sum process (Theorem 3.1), posterior characterizations (Theorem 4.1 and Propositions 4.1–4.3), and a three-component predictive rule for new samples (Proposition 4.4). The authors propose Bayesian diversity measures and an unseen-species entropy (Sections 2.2 and 4.6), and they compare generalized-gamma and gamma versions of the model on simulated data and on a 12-sample microbiome count dataset (Section 6). The main theoretical derivations appear internally consistent and are developed through standard Poisson-process and subordinator calculus, with proofs in Appendix A; the main substantive problem is the claimed distinction between structural and sampling zeros, which the fitted model does not actually implement.","tokens_in":35230,"tokens_out":7538,"duration_ms":79745,"significance":"If the mathematical claims are correct, the paper provides a useful unifying treatment of hierarchical Poisson IBP models, with explicit joint distributions, exact marginal sampling, and posterior samplers, extending earlier work of James and others in the species-sampling literature. The compound-Poisson representation and the Gibbs-partition interpretation in Proposition 4.1 are genuine technical contributions, and the GG-specific formulas connecting to Stirling numbers are useful. The empirical part is suggestive but limited: it compares only two versions of the same model, and the advertised practical contribution of distinguishing technical from biological zeros is not supported by the model's posterior structure. The paper's potential significance therefore rests on the theoretical development rather than on the microbiome-specific zero-inflation claims, which currently need correction.","major_comments":[{"comment":"The advertised distinction between structural zeros and sampling zeros is not realized by the fitted gamma/GG PHIBP. For any species observed in at least one group, H_l > 0 almost surely, and Proposition 4.2 gives the posterior local rate as σ̃_{j,l}(H_l) = σ̂_{j,l}(H_l) + Σ_{k=1}^{X_{j,l}} S_{j,k,l}. For the gamma and generalized gamma Levy densities used in Section 6, σ̂_{j,l}(H_l) is the value of an infinite-activity subordinator after an Esscher transform; in the gamma case it is Gamma(θ_j H_l, ζ_j + Σ_i γ_{i,j}), which has no atom at zero. Hence Pr(σ̃_{j,l}(H_l)=0 | N_{j,l}=0) = 0. A zero count for a globally observed species is always a sampling zero under the implemented model, despite Section 2.1 claiming that the posterior 'reflects whether a zero count Nj,l = 0 likely corresponds to a structural zero (where the posterior for the latent rate σj,l becomes concentrated at zero)'. The posterior can be small, but it cannot represent biological absence, so the structural-versus-sampling-zero decomposition and the diversity/unseen-species interpretations built on it in Sections 2.2 and 4.6 are not supported by the model actually fitted. The authors should either add an explicit zero-inflation layer that places prior mass at zero or substantially reframe these claims.","section":"Section 2.1 and Proposition 4.2, Eq. (4.5)"},{"comment":"The empirical evaluation does not test the zero-handling claim. The experiments compare GG PHIBP against Gamma PHIBP, but neither prior places any mass at σ_{j,l}=0, so differences in frequency-of-frequency distributions, test log-likelihoods, and diversity measures cannot validate the structural-zero mechanism advertised in the abstract and Section 2.1. A comparison against a model with an explicit zero-inflation component, or a direct posterior check for concentration near zero versus positive mass, would be needed to support the claim that the framework 'explicitly distinguish[es] between technical and biological zeros'. As it stands, the microbiome experiments only show that the generalized-gamma subordinator fits the power-law frequency-of-frequency pattern better than the gamma subordinator, which is a weaker and more conventional claim.","section":"Section 6.2 and Appendix B"},{"comment":"The unseen-species entropy U_{j,M_j+1} is presented as a novel contribution, but its identification rests on the same problematic zero decomposition. The quantity is defined on the posterior rates of completely new species, and Proposition 4.5 provides a sampling algorithm for its scaled version. However, the accompanying discussion in Section 4.6 connects this quantity to the number of previously unobserved species Q2 and to the paper's broader claims about handling sparse co-occurrence. Because the model has no atom at zero, the 'unseen' component can only describe low-rate species, not species that are truly absent from the community. The text should either clarify that the model addresses rarity rather than absence or add a mechanism that can represent genuine structural zeros.","section":"Section 4.6 and Proposition 4.5"}],"minor_comments":[{"comment":"The proof of Theorem 4.1 refers to 'Proposition 5.2' where Proposition 4.2 is meant, and the same cross-reference error appears in Remark A.3.","section":"Appendix A and Remark A.3"},{"comment":"Equation (3.7) is mis-rendered: as printed, the displayed expression is not a well-formed joint density, and the notation ψ^{(c_{j,k,l})}_j is introduced only later in the text. The formula should be rewritten or the notation defined immediately.","section":"Section 3.4, Eq. (3.7)"},{"comment":"The sentence 'We next treat the case of Theorem 3.3' appears to refer to Theorem 3.1, and there are several typos such as 'Propostion' and 'Thoerem' in the appendix; these should be corrected before publication.","section":"Appendix A, proof of Theorem 3.1"},{"comment":"The text states that the framework is accompanied by 'freely available resources to validate our model's performance', but no repository or URL is provided in the manuscript; a link or reference to public code would improve reproducibility.","section":"Section 1 and Section 6"}],"recommendation":"major_revision","confidential_remarks":"The theoretical core of the paper appears sound and is likely publishable after the zero-decomposition claims are fixed or substantially reframed. I would not recommend acceptance in the current form because the abstract and Section 2 advertise a structural-versus-sampling-zero capability that the gamma/GG PHIBP does not possess. The heavy reliance on the authors' own unpublished preprint [18] is transparent, and the core results do not appear circular. The empirical section is thin but acceptable as an illustration once the over-claims are removed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on the PHIBP paper. The theoretical machinery is genuinely useful: a three-level CRM construction where the base measure of a Poisson IBP is itself a CRM, plus compound Poisson representations and an Allocation process that give exact sampling and posterior/predictive rules. The Gibbs-partition connections and explicit forms for gamma and generalized gamma cases are real contributions, and the main theorems seem correct. The paper is honest that the generative process is assembled from known components, which is fine.\n\nThe soft spot is the zero-inflation interpretation. Section 2.1 promises that the model can classify a zero count as structural (posterior rate concentrated at zero) versus sampling (mass at positive values). But for any species observed anywhere, the posterior local rate is a continuous random variable on (0,∞) — in the gamma case it's Gamma(θ_j H_l + n_{j,l}, ...), which has zero mass at zero. So the model cannot actually represent biological absence. The paper's own Section 2.2 says the posterior is 'a full distribution, not a point mass at zero,' which contradicts the earlier framing. This doesn't break the math, but it undercuts the diversity and unseen-species interpretations that motivate the application.\n\nThe empirical section is weaker. There's no code or data release, no baseline comparisons against ecological or other count models; the real-data experiment is a within-model comparison of GG versus gamma priors. That's enough to show the GG prior is more flexible, but not enough to validate the PHIBP as a microbiome tool. There are also typos and cross-reference errors (e.g., Proposition 5.2 should be 4.2, equation 3.7 is garbled) that a careful copyedit would fix.\n\nOverall, this is a sound theory paper with an overstated applications story. Worth refereeing seriously, but the authors need to either drop the structural-zero language or add a genuine zero-inflation component, and they should release code and data if they want the microbiome claims taken at face value. I'd send it to review.","headline":"Sound and original theory for hierarchical Poisson IBPs, but the advertised structural-zero interpretation is not realized by the fitted gamma/GG models, and the empirical validation is too thin to support the microbiome claims.","tokens_in":35704,"tokens_out":2221,"would_cite":true,"duration_ms":23124,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60G09","62F15","60G57","62P10","60C05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the Poisson Hierarchical Indian Buffet Process reduces the joint marginal distribution of sparse grouped counts to an exact compound Poisson representation, enabling tractable posterior and predictive inference for…","keywords":["Bayesian nonparametrics","Hierarchical Indian Buffet Process","species sampling models","microbiome count data","zero-inflation","unseen species","compound Poisson process","generalized gamma process"],"falsifier":"Simulate microbiome-like data with a genuine extra zero-inflation layer, e.g., each observed count independently set to zero with probability p beyond the Poisson rate, then fit both the PHIBP and a PHIBP extended with such a zero-inflation component; if the extended model recovers p > 0 with substantially better predictive likelihood, the PHIBP's zero decomposition is not identified.","tokens_in":34629,"feed_emoji":"🦠","tokens_out":5004,"duration_ms":51017,"temperature":0.7,"pith_summary":"This paper introduces the Poisson Hierarchical Indian Buffet Process (PHIBP), a Bayesian nonparametric prior for sparse, grouped count data in which an infinite catalogue of species is shared across groups through a hierarchy of Poisson random measures. It establishes that the joint marginal distribution of the group-sum counts is exactly a compound Poisson process driven by a latent allocation process, giving a tractable exact sampling scheme and explicit posterior and prediction rules. The authors argue this provides a principled treatment of zeros in microbiome count matrices, distinguishing species that are truly absent from species that are present but undetected, and yields model-based diversity and unseen-species estimates. The practical claim is that a generalized gamma specification captures rare species better than a gamma prior, avoiding the artifacts of pseudocount-based methods.","feed_headline":"New hierarchical process samples sparse species counts exactly","feed_subtitle":"A Poisson Indian buffet prior turns microbiome zeros and unseen species into tractable Bayesian quantities.","key_machinery":"The species allocation process $A_J = (\\sum_{l=1}^\\infty \\xi_{j,l}\\,\\delta_{Y_l}: j\\in[J])$, where $\\xi_{j,l}$ are Poisson counts of latent OTUs for species $Y_l$ in group $j$, is the hidden combinatorial engine of the paper. It converts the three-level hierarchy of Poisson random measures into a multivariate Poisson Indian buffet process, whose thinning leads to compound Poisson representations using zero-truncated Poisson (tP) and mixed truncated Poisson (MtP) variables, with Multinomial allocation of counts across groups. Finite Gibbs exchangeable partition functions $\\Xi^{[n_{j,l}]}_{x_{j,l}}$ carry the posterior normalization, and the generalized gamma versus gamma choice of Lévy densities controls whether rare species are preserved in the frequency-of-frequency behavior.","core_discovery":"The paper's central discovery is that the PHIBP marginal distribution, posterior, and predictive rules reduce to compound Poisson representations built from components already in the literature. Theorem 3.1 shows the sum process is distributionally equal to a compound Poisson process whose atoms are species with MtP-distributed latent OTU counts, with group allocation following a Multinomial distribution, enabling exact generative sampling of the marginal. Theorem 4.1 gives the posterior as a decomposition in which observed species have posterior rates split into contributions from OTUs not yet seen in the sample and contributions from observed OTUs, while unseen species retain a residual completely random measure. Proposition 4.4 decomposes prediction for a new sample into arrivals of completely new species, new OTUs of existing species, and additional counts for known OTUs, which together answer classical unseen-species questions in a form suited to sequencing data where the total number of reads is itself random.","pith_inferences":["Because the construction uses only Poisson random measures and zero-truncated Poisson variables, the same compound-Poisson machinery should transfer to other sparse count domains such as genetic variant discovery or text token frequencies without new theory.","When exact sequence variants are observed, the latent OTU counts can be treated as fixed inputs, turning the model into a direct hierarchical sampler for strain-level fitness rates; this is an extension the paper outlines but does not develop empirically.","The predictive unseen entropy $U_{j,M_{j+1}}$ could serve as a sampling-design criterion, indicating which group is expected to yield the most diversity among species not yet seen in any sample.","One could test the zero-handling claim by fitting the PHIBP alongside a version with an explicit extra zero-inflation layer on synthetic data generated with known dropout; the paper does not include such a comparison."],"forward_implications":["Theorem 3.1 gives an exact generative sampler for the PHIBP marginal distribution, so the complex multi-dimensional count distributions can be simulated directly without approximating the infinite hierarchy.","The posterior representation splits each observed species' latent rate into an unobserved-OTU part and an observed-OTU part, which is what lets the model express uncertainty about zeros rather than collapsing them to zero.","Proposition 4.4 divides prediction into completely new species, new OTUs of known species, and additional counts for known OTUs, yielding answers to unseen-species questions under random total reads.","On the real microbiome dataset, the generalized gamma PHIBP matches the power-law frequency-of-frequency distribution of the test data and gives higher predictive likelihood than the gamma PHIBP, which underestimates rare species and inflates beta diversity.","The same framework extends to covariates, technical variation adjustments, and strain-trait modeling, with the latent OTU components providing interpretable fitness rates."],"supporting_citations":[{"why":"Supplies the Poisson Indian buffet process calculus and multivariate IBP machinery that the PHIBP decomposition builds on.","marker":"[16]"},{"why":"Provides the discretization process and joint distributions of species rates and counts that underlie the compound Poisson representations.","marker":"[32]"},{"why":"Gives posterior representations for normalized random measures and the gamma-case EPPF used in the posterior and diversity results.","marker":"[19]"},{"why":"Supplies random mapping and occupancy results used for the finite Gibbs partition structure in Proposition 4.1.","marker":"[23]"},{"why":"Provides the EPPF framework and partition identities that connect the normalization constants to Stirling numbers and Gibbs partitions.","marker":"[34]"},{"why":"Introduces the infinite gamma-Poisson feature model whose Poisson IBP construction the PHIBP extends hierarchically.","marker":"[47]"},{"why":"Contributes frequency-of-frequency distributions and mixed truncated Poisson variables used in the marginal and prediction descriptions.","marker":"[51]"},{"why":"Provides the real microbiome dataset with 42% zeros that motivates and validates the zero-handling and diversity claims.","marker":"[49]"}],"fun_headline_variants":["PHIBP: exact sampling for sparse microbiome counts","Compound Poisson view solves unseen species in microbiome data","New Bayesian process borrows strength across microbiome groups","Exact generative sampling for infinite-species count models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes that every zero in the observed count matrix is fully explained by the latent Poisson sampling rate, so there is no separate zero-inflation mechanism (such as PCR dropout) acting on counts; if such a mechanism exists, the split between structural and sampling zeros is not identified.","fun_headline_variants_meta":{"raw":{"variants":["PHIBP: exact sampling for sparse microbiome counts","Compound Poisson view solves unseen species in microbiome data","New Bayesian process borrows strength across microbiome groups","Exact generative sampling for infinite-species count models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000499,"raw_usage":{"total_tokens":2438,"prompt_tokens":932,"completion_tokens":1506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":1445}},"tokens_in":548,"tokens_out":1506,"duration_ms":10926,"temperature":1.0,"reasoning_tokens":1445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:59:42.006573+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate microbiome-like data with a genuine extra zero-inflation layer, e.g., each observed count independently set to zero with probability p beyond the Poisson rate, then fit both the PHIBP and a PHIBP extended with such a zero-inflation component; if the extended model recovers p > 0 with substantially better predictive likelihood, the PHIBP's zero decomposition is not identified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Poisson Indian buffet process calculus and multivariate IBP machinery that the PHIBP decomposition builds on."},{"cited_title":"[1997], ‘Partition structures derived from Brownian motion and stable subordinators’, Bernoulli 3, 79–96","cited_arxiv_id":null,"evidence_quote":"Provides the discretization process and joint distributions of species rates and counts that underlie the compound Poisson representations."},{"cited_title":"F., Lijoi, A","cited_arxiv_id":null,"evidence_quote":"Gives posterior representations for normalized random measures and the gamma-case EPPF used in the posterior and diversity results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies random mapping and occupancy results used for the finite Gibbs partition structure in Proposition 4.1."},{"cited_title":"[2006], Combinatorial stochastic processes, Lectures from the 32nd Summer School on Probability Theory held in Saint-Flour, July 7–24, 2002","cited_arxiv_id":null,"evidence_quote":"Provides the EPPF framework and partition identities that connect the normalization constants to Stirling numbers and Gibbs partitions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the infinite gamma-Poisson feature model whose Poisson IBP construction the PHIBP extends hierarchically."},{"cited_title":"and Walker, S","cited_arxiv_id":null,"evidence_quote":"Contributes frequency-of-frequency distributions and mixed truncated Poisson variables used in the marginal and prediction descriptions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the real microbiome dataset with 42% zeros that motivates and validates the zero-handling and diversity claims."}],"review_version":1}