REVIEW 1 major objections 5 minor 32 references
Extended feature allocation models
T0 review · 1 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that, within extended feature allocation models, the predictive distribution for new features depends only on sample size exactly when the prior point process is Poisson, and only on sample size plus the number of…
desk verdict The Palm-calculus framework and the DPP prior are genuinely useful, but the sufficientness theorems are over-stated: the 'only if' directions rest on an unproved identifiability step and, as written, appear to be false. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the reduced Palm kernel of the prior point process $\Psi = \sum_j \delta_{(X_j,S_j)}$ on $X \times (0,1]$. For conditioning points $(x,s)$, the reduced Palm version $\Psi^!_{x,s}$ is the law of the remaining atoms after forcing atoms at $x$ with marks $s$ and removing them; it enters Theorem 2's posterior representation and, through thinning, Theorem 3's predictive law for $Z'_{n+1}$. Lemma 1 and Lemma 2 are the identities that carry the argument: the reduced Palm law is location-independent if and only if $\Psi$ is Poisson, and it depends only on the number of conditioning points if and only if $\Psi$ is mixed Poisson or mixed binomial. Everything else in the Bayesian analysis is a Palm-calculus computation: the factorial moment disintegration, the Campbell-Little-Mecke formula, and the Laplace functional ratio that defines the predictive law.
What would settle it
On a finite label space, enumerate possible reduced Palm kernels for a two-point prior and compute the induced predictive law, the Bernoulli-process thinning in equation (7), for both a Poisson and a non-Poisson kernel with the same factorial moments; if two different kernels yield the same predictive law for every sample, the only-if direction of Theorem 4 fails. Even a single analytic pair of non-Poisson processes with identical new-feature predictive laws would refute the stated characterization.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a pair of characterization theorems. Theorem 4 states that, within model (3), the distribution of the new-feature process $Z'_{n+1}$ from Theorem 3 is a function of the sample size $n$ alone if and only if the underlying point process $\Psi$ is Poisson. Theorem 5 states that this distribution is a function of $n$ and the number of distinct features $k$ alone if and only if $\Psi$ is a mixed Poisson or mixed binomial process. The proof route that supplies these results is a characterization of the reduced Palm kernel: a point process whose reduced Palm versions $\Psi^!_{x,s}$ have a law independent of the conditioning points must be Poisson, and a process whose reduced Palm laws depend only on the number of conditioning points must be mixed Poisson or mixed binomial. Thus the simplicity of predictive distributions is exactly the simplicity of Palm kernels, and the paper derives a characterization of the Poisson process as a byproduct.
Load-bearing premise
The only-if directions of both theorems assume that the distribution of the thinned new-feature process $Z'_{n+1}$ uniquely determines the reduced Palm kernel of $\Psi$; the proof states that this is clear rather than proving that the map from Palm kernels to predictive laws is injective.
Editorial extensions
If this is right
- Under a completely random measure (Poisson) prior, no predictive statement about new features can use the observed labels or the frequency spectrum; the entire new-feature law is fixed by the sample size.
- To obtain predictions that depend on both sample size and number of distinct features, it is necessary and sufficient to use a mixed Poisson or mixed binomial prior, and stable beta scaled processes and product-form feature models are recovered as special cases.
- Any extended feature prior outside these classes yields new-feature predictions that depend on labels or frequencies, giving a principled reason to choose such priors for the unseen-feature problem.
- The reduced-Palm characterization is a standalone result: a point process whose reduced Palm law does not depend on the conditioning point is Poisson.
- The independently marked determinantal point process prior makes the predictive locations of new features depend on the observed locations, which is what allows the forest application to locate unseen trees.
Reading between the lines
- If the identifiability step that the supplementary proof marks as clear fails, the only-if directions of Theorems 4 and 5 would collapse to one-way statements; the theorems should then be reformulated as characterizing the simplest predictive laws rather than the priors themselves.
- A natural cross-check is to test the forest application against a Poisson prior with matching mean measure: the repulsive prior should concentrate posterior predictive mass closer to the observed configuration, and a formal calibration of coverage over repeated surveys would quantify the gain.
- The sufficientness postulates constrain only the new-feature component $Z'_{n+1}$; they say nothing about the probability of re-observing old features, so prior selection for joint prediction of old and new features would need additional predictive functionals.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces extended feature allocation models, in which the latent feature process Ψ is an arbitrary simple point process on X×(0,1] and observations are i.i.d. Bernoulli processes directed by the random measure μ(B)=∫ s 1_B(x)Ψ(dx ds). The main theoretical contributions are: (i) a general Palm-calculus treatment giving the marginal, posterior, and predictive laws (Theorems 1-3); (ii) two sufficientness postulates — the new-feature predictive law depends on the sample only through n iff Ψ is Poisson (Theorem 4), and only through n and k iff Ψ is mixed Poisson or mixed binomial (Theorem 5); and (iii) specialization to Poisson, mixed Poisson, mixed binomial, and independently marked determinantal point process priors, with an application to forest surveys.
Significance. Should the characterizations be fully established, Theorems 4 and 5 are substantial: they unify previous sufficientness results for CRMs, scaled stable beta processes, and product-form EFPMs within a single point-process framework, and they give practitioners a clear criterion for prior selection. The Palm-calculus machinery in Theorems 1-3 is clean, and the specializations to Poisson, mixed Poisson, and mixed binomial priors are consistent with known results. The paper also offers a novel characterization of the Poisson process via its reduced Palm kernel and a useful DPP-based spatial model. However, the only-if directions of the two central theorems rest on an unproved identifiability step, so the main characterization claim is not yet fully demonstrated.
major comments (1)
- [Supplementary S5.1 (proofs of Theorems 4 and 5)] The only-if directions of Theorems 4 and 5 are not proved as written. The proof says "it is clear that the predictive distribution of Z'_{n+1} depends on the sampling information as the law of μ', or equivalently Ψ', in Theorem 2 does," but this hides a nontrivial inversion. The observable object Z'_{n+1} is a Bernoulli process directed by μ', and μ' is a mixture over S* of tilted reduced Palm versions of Ψ (Theorem 2(i)-(ii) and Eq. (7)). To conclude from constancy of the predictive law in (x*,m) that each reduced Palm kernel Ψ!_{x*,s*} is independent of (x*,s*), one needs three inversions: (a) the law of a Cox process determines the law of its directing random measure; (b) the law of μ' determines the law of the marked point process Ψ'; and (c) the mixture over S* with weights proportional to s*^m(1-s*)^{n-m}ρ^{(k)}(ds*|x*) E[e^{∫ n log(1-t)Ψ!_{x*,s*}}] determines the family of reduced Palm kernels. Part (a) is standard but not stated; part (c) is nontrivial and no argument is provided. Without (c), different reduced Palm kernels could in principle produce the same mixture predictive law for all observed x*,m, so Lemma 1 cannot be invoked. The same gap affects Theorem 5, whose proof says only "arguing as in the proof of Theorem 4." This is a load-bearing point for the paper's central characterization result.
minor comments (5)
- [Section 3.1, Theorem 1] The displayed formula is a density with respect to \tilde m^{(k)}_ξ(dx*), not a probability itself; the wording "The probability ... is" should be "the probability measure has density" or similar.
- [Section 7, last paragraph] The claim that "for any Poisson, mixed Poisson, or mixed binomial prior, the probability of re-observing a feature depends exclusively on the sample size and the frequency of that feature" is too strong when the kernel ρ(ds|x) depends on the label x; in Corollaries 1-3 the marginal density of S*_ℓ is proportional to s^{m_ℓ}(1-s)^{n-m_ℓ}ρ(ds|x*_ℓ), so it depends on x*_ℓ unless ρ is label-independent.
- [Supplementary S5.2, proof of Lemma 1] The last step of the proof is terse: after showing that each Φ!_{x} is Poisson, the conclusion that Φ itself is Poisson invokes the fact that the family of Palm distributions, together with the mean measure, characterizes the law of a simple point process; this should be stated explicitly with a precise reference.
- [Section 6.1] The sentence "the infinitesimal probability that an unobserved tree would occupy position dx equals E{Ψ′(dx × (0,1])}" uses "probability" where the object is a mean measure; please rephrase to avoid confusion with a density.
- [Section 2.3] The claim that "extended feature allocation models induce the entire class of regular feature allocation models admitting an efpf" is nontrivial and is stated without proof; a precise reference or a short argument should be supplied.
Circularity Check
No circular reduction: Theorems 4–5 are proved from independent Palm-characterization lemmas; self-citations are contextual and the empirical Bayes fit does not enter the theoretical claims.
full rationale
The central characterization theorems are not circular. Theorem 4 is reduced in Supplementary Section S5.1 to Lemma 1, and Theorem 5 to Lemma 2; both lemmas are stated independently of the predictive setup and proved from classical point-process facts, with Lemma 2 built on Kallenberg (1973) and Lemma 1 on Slivnyak–Mecke and the Palm characterization of point-process laws. The 'if' directions are independently verified by the concrete posterior and predictive computations in Corollaries 1–3, which derive the Poisson, mixed Poisson, and mixed binomial predictive forms from Theorem 2 rather than from the sufficientness postulates. Thus the statements do not assume their conclusions. The paper's use of earlier work by Camerlenghi et al. (2023) and Ghilotti et al. (2024) is for context and for reconciling scaled processes with Theorem 5; it is not the proof of Theorems 4–5. The empirical Bayes selection of DPP hyperparameters in Section 6.1 is standard estimation practice and is confined to the application; it does not feed back into the theoretical derivations. One genuine gap should be flagged, although it is not circularity: Supplementary Section S5.1 asserts 'it is clear' that the predictive law of Z'_{n+1} depends on the sample exactly as the law of mu' or Psi' in Theorem 2 does, and the subsequent proof transfers the sufficientness assumption to a property of the reduced Palm kernel. The needed identifiability step—that the mixture over S* and the Bernoulli thinning can be inverted to recover the Palm-kernel family—is not demonstrated. If false, the 'only if' directions would be incomplete; however, this is an omitted identifiability argument rather than a reduction of the conclusion to its own input, so it does not affect the circularity score.
Assumptions & free parameters
free parameters (2)
- DPP hyperparameters (a, b, rho, alpha) =
Estimated by empirical Bayes, values not reported
- Simulation constants (marks Beta(1,5) or Beta(1,20), rho=100, alpha=0.0535) =
Hand-chosen
assumptions (6)
- standard math Palm calculus and Campbell-Little-Mecke formulas for locally finite point processes
- standard math Kallenberg (1973) Theorem 5.3 characterization of mixed Poisson/binomial processes via Palm kernels
- domain assumption The statistical model (3): Zi | mu iid BeP(mu), mu is a functional of a simple point process Psi, with conditionally i.i.d. Bernoulli indicators A_ij with probabilities S_j > 0
- domain assumption The k-th factorial moment measures of Psi are sigma-finite for all k, and Z_i(X) < infinity a.s.
- ad hoc to paper Identifiability: the law of Z'_n+1 determines the law of mu' and hence the reduced Palm kernel
- ad hoc to paper In the forest application, tree locations follow a Gaussian determinantal point process and detection probabilities are i.i.d. beta
Cite this review
Pith. "Pith review of Extended feature allocation models." pith.science (2026). https://pith.science/paper/4IRJ3MIE
@misc{pith2026250210257,
author = {Pith},
title = {Pith review of: Extended feature allocation models},
year = {2026},
howpublished = {\url{https://pith.science/paper/4IRJ3MIE}},
note = {Machine review of arXiv:2502.10257}
}
read the original abstract
Feature allocation models are Bayesian nonparametric tools tailored to data in which each observation can simultaneously exhibit multiple characteristics, or features. A fundamental limitation of standard formulations is that feature labels are assumed to be independent and identically distributed, and therefore play no role in posterior inference. The present paper introduces a unified Bayesian framework for extended feature allocation models, in which feature labels and proportions are modeled jointly, thereby enabling the simultaneous discovery of features and learning of dependencies among their labels. Building on point process theory, we develop a full Bayesian analysis of these models. Within this general setting, we also characterize previously proposed priors as those leading to poor predictive distributions, which cannot capture label dependencies and are insensitive to the observed frequency spectrum. Our methodology is designed to move beyond such standard formulations by leveraging the information carried by feature labels. We demonstrate the usefulness of our approach by introducing: (i) a Cox process prior that clusters genomic variant embeddings while predicting new variants and new variant clusters; (ii) a determinantal point process prior for repeated forest surveys, where prediction concerns both the number and the locations of unobserved trees.
Figures
Reference graph
Works this paper leans on
-
[1]
Ascolani, F., B. Franzolini, A. Lijoi, and I. Prünster (2024). Nonparametric priors with full- range borrowing of information.Biometrika 111(3), 945–969. Ayed, F. and F. Caron (2021). Nonnegative bayesian nonparametric factor models with com- pletely random measures.Statistics and Computing 31, 1–24. Bacallado, S., M. Battiston, S. Favaro, and L. Trippa (...
work page 2024
-
[2]
This distinction, while subtle, is crucial
However, it is important to note that in Kallenberg’s theorem, a point process is defined as a random boundedly finite counting measure, whereas we consider a point process as a random locally finite counting measure. This distinction, while subtle, is crucial. Working with boundedly finite measures excludes Poisson processes with infinite activity. Speci...
work page 1973
-
[4]
The marginal distribution ofZ is recovered from Theorem 1 as follows. SinceΨ is an independently marked process with ground processξ and mark kernel H, then from (Baccelli et al., 2020, Proposition 3.2.14), the reduced Palm versionΨ! x,s is still an independently marked process, with ground process ξ! x and mark kernel H, thus it does not depend on s. By ...
work page 2020
-
[7]
Cambridge University Press. Lavancier, F., J. Møller, and E. Rubak (2015). Determinantal point process models and sta- tistical inference. Journal of the Royal Statistical Society: Series B (Statistical Methodol- ogy) 77(4), 853–877. Lavancier, F. and E. Rubak (2023, October). On simulation of continuous determinantal point processes. Statistics and Compu...
work page 2015
-
[12]
Springer. Kallenberg, O. (2021).Foundations of Modern Probability. Springer. Last, G. and M. Penrose (2017). Lectures on the Poisson process, Volume
work page 2021
-
[18]
MIT Press. Ghilotti, L., F. Camerlenghi, and T. Rigon (2024). Bayesian analysis of product feature allo- cation models. arXiv preprint arXiv:2408.15806. Gnedin, A. and J. Pitman (2005). Exchangeable Gibbs partitions and Stirling triangles.Zap. Nauchn. Sem. S.-Peterburg. Otdel. Mat. Inst. Steklov. (POMI) 325, 83–102. 34 Good, I.J.andG.H.Toulmin(1956). Then...
arXiv 2024
-
[19]
Thibaux, R. and M. I. Jordan (2007). Hierarchical beta processes and the indian buffet process. In AISTATS, Volume 2, pp. 564–571. Titsias, M. (2007). The infinite gamma-poisson feature model. In J. Platt, D. Koller, Y. Singer, and S. Roweis (Eds.),Advances in Neural Information Processing Systems, Volume
work page 2007
-
[20]
Cur- ran Associates, Inc. 38 Xie, F. and Y. Xu (2019). Bayesian repulsive gaussian mixture model.Journal of the American Statistical Association, 187–203. Zabell, S. L. (1982). W. E. Johnson’s "Sufficientness" Postulate.The Annals of Statistics 10(4), 1090 –
work page 2019
Show all 32 references
-
[22]
Møller, J
Curran Associates, Inc. Møller, J. and R. P. Waagepetersen (2003).Statistical inference and simulation for spatial point processes. CRC Press. Muliere, P. and S. Walker (1999). A characterization of a neutral to the right prior via an extension of Johnson’s sufficientness post...
2003
-
[23]
is equivalent to R X×(0,1] sν(dx ds) < ∞ (Proposition S2)
Indeed, when Ψ is a Poisson process with mean measureν, the condition Zi(X) < ∞ a.s. is equivalent to R X×(0,1] sν(dx ds) < ∞ (Proposition S2). Therefore, for anyϵ >0, ν(X × (ϵ, 1]) = Z X×(ϵ,1] ν(dx ds) ≤ ϵ−1 Z X×(ϵ,1] sν(dx ds) < ∞. Thus, it follows thatν is σ-finite, as well...
2021
-
[24]
Let LΦ be the Laplace functional of Φ. By (Baccelli et al., 2020, Proposition 3.2.1), for any measurable functions f, g: X → R+, the Palm distribution of a point process satisfies ∂ ∂t LΦ(f + tg) t=0 = − Z X g(x)LΦx(f )MΦ(dx). Since Φ is a mixed binomial processMB(ν, qM ), its...
2020
-
[25]
Pitman, J
Curran Associates, Inc. Pitman, J. (1995). Exchangeable and partially exchangeable random partitions. Probability theory and related fields 102(2), 145–158. Pitman, J. (1996). Some developments of the Blackwell-MacQueen urn scheme. InStatistics, probability and game theory, Vo...
1995
-
[26]
Hence, Φx = δx + Φ! x where Φ! x is distributed as in the statement
E(M ) E e−f (X) k # MΦ(dx). Hence, Φx = δx + Φ! x where Φ! x is distributed as in the statement. The proof for the Palm distribution of orderk follows by induction, by using the Palm algebra property (Baccelli et al., 2020, Proposition 3.3.9), i.e.,Φ! (x1,x2) d = Φ! x1 ! x2 , ...
2020
-
[27]
Lemma S5
Secondly, we state another key result which we need for the proof, corresponding to (Kallen- berg, 2017, Theorem 3.7). Lemma S5. Let Φ be a point process onX and let Cj ↑ X, Cj ∈ X and relatively compact, j ≥
2017
-
[28]
S4 Proof of Theorems 1 and 2 The proof is based on the study of the Laplace functional of µ
Finally, by an application of Lemma S5, we conclude thatΦ is either a mixed Poisson or mixed binomial process. S4 Proof of Theorems 1 and 2 The proof is based on the study of the Laplace functional of µ. We remind that, for any measurable function f : X → R+, the Laplace funct...
2017
-
[29]
S5.2 Proof of Lemma 1 If Φ is a Poisson process, thenΦ! x d = Φ ! y d = Φ by the multivariate Mecke equation (Last and Penrose, 2017, Theorem 4.4)
Then, the proof follows from Lemma 1 which characterizes the Poisson process as the unique point process for which the reduced Palm kernelΨ! x∗,s∗ does not depend on(x∗, s∗). S5.2 Proof of Lemma 1 If Φ is a Poisson process, thenΦ! x d = Φ ! y d = Φ by the multivariate Mecke eq...
2017
-
[30]
Using this characterization, proceed as follows to prove Lemma 2
S5.4 Proof of Lemma 2 The proof of our lemma requires the use of Lemma S3. Using this characterization, proceed as follows to prove Lemma 2 . First, (ii) implies (i) by using (Kallenberg, 1973, Theorem 5.3). Indeed, by settingk = 1, condition (ii) states thatΦ! x does not depe...
1973
-
[31]
Moreover, since Ψ is a Poisson process, by Lemma 1, it holdsΨ! x∗,s d = Ψ
In particular, we first observe that˜m(k) ξ (dx∗) = Qk ℓ=1 G0(dx∗ ℓ ) and ρ(k)(ds | x∗) = Qk ℓ=1 ρ(dsℓ | x∗ ℓ ). Moreover, since Ψ is a Poisson process, by Lemma 1, it holdsΨ! x∗,s d = Ψ. Consequently, the expected value in the marginal expression of Theorem 1 equals E n e R X...
2017
-
[32]
, S∗ k) and the law ofµ′, conditionally to S∗
For the posterior distribution ofµ, expressed in Theorem 2, we need to determine the law of the vectorS∗ = (S∗ 1 , . . . , S∗ k) and the law ofµ′, conditionally to S∗. From point (i) of Theorem 2 and (S16), which does not depend ons, the S∗ ℓ’s are independent with marginal la...
2020
-
[440]
James, L. F. (2017, October). Bayesian Poisson calculus for latent feature modeling via gener- alized Indian Buffet Process priors.The Annals of Statistics 45(5), 2016–2045. James, L. F., J. Lee, and A. Pandey (2023). Bayesian analysis of generalized hierarchical Indian buffet...
2023 arXiv
-
[449]
De Iorio, and J
Franzolini, B., M. De Iorio, and J. Eriksson (2023). Conditional partial exchangeability: a probabilistic framework for multi-view clustering.arXiv preprint arXiv:2307.01152. Ghahramani, Z. and T. Griffiths (2005). Infinite latent feature models and the indian buffet process. ...
2023 arXiv
-
[500]
Favaro, and L
Bacallado, S., S. Favaro, and L. Trippa (2013). Bayesian nonparametric analysis of reversible markov chains.The Annals of Statistics, 870–896. Baccelli, F., B. Błaszczyszyn, and M. Karray (2020). Random measures, point processes, and stochastic geometry.HAL preprint available ...
2013
-
[599]
Ord, K. (1978). How many trees in a forest?The Mathematical Scientist 3, 23–33. Orlitsky, A., A. T. Suresh, and Y. Wu (2016). Optimal prediction of the number of unseen species. Proceedings of the National Academy of Sciences 113(47), 13283–13288. Petralia, F., V. Rao, and D. ...
1978
-
[836]
Broderick, T., A. C. Wilson, and M. I. Jordan (2018). Posteriors, conjugacy, and exponential families for completely random measures.Bernoulli 24(4B), 3181 –
2018
-
[900]
Prünster, I. (2002). Random probability measures derived from increasing additive processes and their application to Bayesian statistics. Ph. D. thesis, University of Pavia. Regazzini, E. (1978). Intorno ad alcune questioni relative alla definizione del premio secondo la teori...
2002
-
[1099]
Bayesian calculus and predictive characterizations of extended feature allocation models
Zabell, S. L. (1995). Characterizing markov exchangeable sequences. Journal of Theoretical Probability 8, 175–178. Zabell, S.L.(2005). Symmetry and its discontents. CambridgeStudiesinProbability, Induction, and Decision Theory. Cambridge University Press, New York. Essays on t...
1995
-
[1283]
Rediscoveryofgood–turingestimatorsviabayesian nonparametrics
Favaro, S., B.Nipoti, andY.W.Teh(2016). Rediscoveryofgood–turingestimatorsviabayesian nonparametrics. Biometrics 72(1), 136–145. Fong, E., C. Holmes, and S. G. Walker (2023). Martingale posterior distributions.J. R. Stat. Soc. Ser. B. Stat. Methodol. 85(5), 1357–1391. Fortini,...
2016
-
[1291]
Prünster, and S
Lijoi, A., I. Prünster, and S. G. Walker (2008). Investigating nonparametric priors with gibbs structure. Statistica Sinica, 1653–1668. Lo, A. Y. (1991). A characterization of the dirichlet process. Statistics & Probability Let- ters 12(3), 185–187. Lui, A., J. Lee, P. F. Thal...
2008 arXiv
-
[1448]
Argiento, F
32 Beraha, M., R. Argiento, F. Camerlenghi, and A. Guglielmi (2023). Bayesian mixtures models with repulsive and attractive atoms.arXiv preprint arXiv:2302.09034. Beraha, M. and J. E. Griffin (2023). Normalised latent measure factor models.Journal of the Royal Statistical Soci...
2023 arXiv
-
[1799]
James, L. F. (2006). Poisson calculus for spatial neutral to the right processes.The Annals of Statistics 34(1), 416 –
2006
-
[2324]
Hough, J. B., M. Krishnapur, Y. Peres, and B. Viràg (2006). Determinantal processes and independence. Probability Surveys 3, 206–229. Hu, Y., K. Zhai, S. Williamson, and J. Boyd-Graber (2012). Modeling images using transformed Indian buffet processes. InInternational Conferenc...
2006 arXiv
-
[3221]
Sufficientness
Camerlenghi, F. and S. Favaro (2021). On Johnson’s “Sufficientness” postulates for feature- sampling models. Mathematics 9(22). Camerlenghi, F., S. Favaro, L. Masoero, and T. Broderick (2023). Scaled process priors for Bayesian nonparametric estimation of the unseen genetic va...
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.