REVIEW 2 major objections 1 minor 35 references
Bayesian Mixture Models for Histograms: with Applications to Large Datasets
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Bayesian mixture models recover underlying distributions from histogram bin counts alone via reversible jump MCMC.
desk verdict The paper puts reversible-jump MCMC on binned counts together with a Dirichlet-process layer for clustering multiple histograms; the identifiability of the mixtures from coarse bins is the main open question. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Reversible jump MCMC applied to normal mixture models on histogram bin counts, with a Dirichlet process prior for simultaneous clustering of multiple histograms.
What would settle it
A simulation study in which the same dataset is available both in raw form and as histograms; the mixture recovered from the binned version should closely match the mixture fitted directly to the raw observations.
Extended reading notes
Core claim
The central claim is that placing a prior on the number of mixture components and performing reversible jump MCMC allows Bayesian inference of the underlying continuous distribution from binned data alone. The framework is extended to multiple histograms by using a Dirichlet process to cluster them, enabling information sharing across populations and providing a posterior probability to assess homogeneity between groups. Some theoretical results support the performance of the methodology.
Load-bearing premise
The data-generating process is well approximated by a mixture of normal distributions whose parameters can be recovered from the observed bin counts.
Editorial extensions
If this is right
- The method recovers population distributions accurately from aggregated data without access to individual observations.
- It scales to large datasets where processing raw records would be computationally prohibitive.
- Multiple histograms can be clustered to share statistical strength and produce posterior probabilities of similarity between populations.
- The modeling framework is stated to be flexible for extension to mixture families other than normals.
- Theoretical support is provided for the consistency and performance of the binned-data inference.
Reading between the lines
- This approach could enable statistical analysis in privacy-restricted settings that release only histograms.
- It suggests a route to density estimation when data arrives already summarized or streamed in aggregate form.
- The clustering extension might apply to comparing distributions across sites or time periods without pooling raw records.
- Direct comparison of binned versus raw inference on benchmark datasets would test how much information is lost in aggregation.
- keywords:[
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces a Bayesian nonparametric approach to inferring underlying distributions from histogram (binned) data by fitting finite or countably infinite mixtures of normals via reversible-jump MCMC, with an extension to simultaneous modeling and clustering of multiple histograms using a Dirichlet process; it claims strong empirical performance on large-scale data together with supporting theoretical results.
Significance. If the recovery of mixture parameters from binned counts proves stable, the method would offer a principled way to perform inference and clustering on aggregated or privacy-protected data, extending reversible-jump and Dirichlet-process techniques to histogram likelihoods.
major comments (2)
- [Theoretical results] Theoretical results section: the manuscript states that theoretical results support performance, yet provides no explicit argument or bound establishing identifiability or posterior consistency when bin widths are comparable to component scales or when components overlap; this directly underpins the claim that parameters can be recovered from binned counts alone.
- [Simulation studies] Simulation or application studies: the reported strong performance on large-scale data is not accompanied by recovery experiments that vary bin width relative to component variance or that quantify posterior multimodality under overlapping components; without such checks the central empirical claim remains unverified.
minor comments (1)
- [Abstract] The abstract does not specify the form of the prior on the number of components (e.g., Poisson or stick-breaking) or the precise reversible-jump proposal mechanism.
Simulated Author's Rebuttal
We thank the referee for the thoughtful and constructive report. The comments identify important gaps in the theoretical justification and empirical validation of parameter recovery from binned data. We address each point below and will revise the manuscript to strengthen these aspects.
read point-by-point responses
-
Referee: [Theoretical results] Theoretical results section: the manuscript states that theoretical results support performance, yet provides no explicit argument or bound establishing identifiability or posterior consistency when bin widths are comparable to component scales or when components overlap; this directly underpins the claim that parameters can be recovered from binned counts alone.
Authors: The manuscript presents theoretical results on posterior consistency for the histogram likelihood under the reversible-jump MCMC scheme for normal mixtures, relying on standard conditions for Dirichlet process mixtures and the continuity of the binned likelihood. We acknowledge that these results do not include explicit identifiability bounds or consistency rates for the regime in which bin widths approach component scales or when components overlap substantially. We will revise the theoretical section to clarify the scope of the existing arguments and add a discussion of the additional conditions required in the overlapping or coarse-binning cases, together with references to related identifiability results for binned mixtures. revision: yes
-
Referee: [Simulation studies] Simulation or application studies: the reported strong performance on large-scale data is not accompanied by recovery experiments that vary bin width relative to component variance or that quantify posterior multimodality under overlapping components; without such checks the central empirical claim remains unverified.
Authors: The current simulation studies focus on large-scale data sets with fixed binning and demonstrate accurate recovery and clustering performance. We agree that systematic variation of bin width relative to component variance and explicit quantification of posterior multimodality under overlap would provide stronger verification of the central claim. We will add a new set of targeted simulation experiments addressing these regimes and report the corresponding recovery metrics and posterior diagnostics in the revised manuscript. revision: yes
Circularity Check
No circularity detected; derivation is self-contained
full rationale
The manuscript introduces a Bayesian nonparametric mixture model for histogram data via reversible-jump MCMC, with a Dirichlet-process extension for multiple histograms. No equations, parameter-fitting steps, or self-citations are presented that reduce any claimed result to a definition or input by construction. The performance claims rest on the standard MCMC posterior rather than any self-referential prediction or renamed empirical pattern. The provided text contains no load-bearing self-citation chains or ansatz smuggling. This is the expected honest outcome for a methodological paper whose central procedure is externally verifiable against simulated or real binned data.
Assumptions & free parameters
assumptions (2)
- domain assumption The underlying density belongs to the class of finite or countably infinite normal mixtures.
- domain assumption Reversible jump MCMC can be implemented to sample from the posterior over both component number and parameters given only bin counts.
Cite this review
Pith. "Pith review of Bayesian Mixture Models for Histograms: with Applications to Large Datasets." pith.science (2026). https://pith.science/paper/HLYQYXSI
@misc{pith2026260624001,
author = {Pith},
title = {Pith review of: Bayesian Mixture Models for Histograms: with Applications to Large Datasets},
year = {2026},
howpublished = {\url{https://pith.science/paper/HLYQYXSI}},
note = {Machine review of arXiv:2606.24001}
}
read the original abstract
In many real-world scenarios, especially those involving privacy constraints or data summarization, data are available only in aggregated forms, such as histograms or frequency tables. This work introduces a novel Bayesian method for inferring the underlying population distribution by fitting a mixture model to binned data. While we focus on mixtures of normal distributions, the framework is flexible and can be extended to other distributional families. We place a prior distribution on the number of mixture components, accommodating both finite and countably infinite mixtures, and perform inference using reversible jump MCMC. The proposed approach demonstrates strong performance on large-scale data, showcasing the potential of nonparametric Bayesian modeling in practical applications. Furthermore, we extend the method to model multiple histograms simultaneously and cluster them using the Dirichlet process. This enables information sharing across populations and provides a principled posterior probability to assess homogeneity between groups. Some theoretical results supporting the performance of our proposed methodology are also discussed.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Alston, C. L. and Mengersen, K. L. (2010). Allowing for the effect of data binning in a B ayesian normal mixture model. Computational Statistics and Data Analysis , 54(4):916--923
2010
-
[2]
Argiento, R., Cremaschi, A., and Vannucci, M. (2020). Hierarchical normalized completely random measures to cluster grouped data. Journal of the American Statistical Association , 115(529):318--333
2020
-
[3]
Beraha, M., Guglielmi, A., and Quintana, F. A. (2021). The semi-hierarchical D irichlet process and its application to clustering homogeneous distributions. Bayesian Analysis , 16(4):1187--1219
2021
-
[4]
and Kim, J
Billard, L. and Kim, J. (2017). Hierarchical clustering for histogram data. Wiley Interdisciplinary Reviews: Computational Statistics , 9(5):e1405
2017
-
[5]
and Dunson, D
Canale, A. and Dunson, D. B. (2011). B ayesian kernel mixtures for counts. Journal of the American Statistical Association , 106(496):1528--1539
2011
-
[6]
Corradin, R., Canale, A., and Nipoti, B. (2021). BNPmix : An R package for B ayesian nonparametric modeling via P itman- Y or mixtures. Journal of Statistical Software , 100(15):1--33
2021
-
[7]
B., Johnson, D
Dahl, D. B., Johnson, D. J., and M \"u ller, P. (2022). Search algorithms and loss functions for B ayesian clustering. Journal of Computational and Graphical Statistics , 31(4):1189--1201
2022
-
[8]
Denti, F., Camerlenghi, F., Guindani, M., and Mira, A. (2023). A common atoms model for the bayesian nonparametric analysis of nested data. Journal of the American Statistical Association , 118(541):405--416
2023
Show all 35 references
-
[9]
Duan, Y., Guo, S., Wang, W., and M \"u ller, P. (2025). Immune profiling among colorectal cancer subtypes using dependent mixture models. Journal of the American Statistical Association , 120(550):671--684
2025
-
[10]
and Denti, F
D’Angelo, L. and Denti, F. (2026). A finite-infinite shared atoms nested model for the bayesian analysis of large grouped data sets. Bayesian Analysis , 21(1):105--138
2026
-
[11]
Escobar, M. D. and West, M. (1995). B ayesian density estimation and inference using mixtures. Journal of the american statistical association , 90(430):577--588
1995
-
[12]
Gau, S., de Dieu Tapsoba, J., and Lee, S. (2014). Bayesian approach for mixture models with grouped data. Computational Statistics , 29(5):1025--1043
2014
-
[13]
and van der Vaart, A
Ghosal, S. and van der Vaart, A. W. (2007). Convergence rates of posterior distributions for non-iid observations. Annals of Statistics , 35(1):192--223
2007
-
[14]
T., Semenova, L., and Rudin, C
Goh, S. T., Semenova, L., and Rudin, C. (2024). Sparse density trees and lists: An interpretable alternative to high-dimensional histograms. INFORMS Journal on Data Science
2024
-
[15]
and Arabie, P
Hubert, L. and Arabie, P. (1985). Comparing partitions. Journal of classification , 2(1):193--218
1985
-
[16]
and James, L
Ishwaran, H. and James, L. F. (2001). Gibbs sampling methods for stick-breaking priors. Journal of the American Statistical Association , 96(453):161--173
2001
-
[17]
and Eilers, P
Lambert, P. and Eilers, P. H. C. (2009). Bayesian density estimation from grouped continuous data. Computational Statistics & Data Analysis , 53(4):1388--1399
2009
-
[18]
Lijoi, A., Pr \"u nster, I., and Rebaudo, G. (2023). Flexible clustering via hidden hierarchical dirichlet priors. Scandinavian Journal of Statistics , 50(1):213--234
2023
-
[19]
Mart \' nez, A. F. and D \' az-Avalos, C. (2024). A model-based approach for clustering binned data. arXiv preprint arXiv:2409.07738
2024
-
[20]
Miller, J. W. and Harrison, M. T. (2018). Mixture models with a prior on the number of components. Journal of the American Statistical Association , 113(521):340--356
2018
-
[21]
L., Kochanek, K
Murphy, S. L., Kochanek, K. D., Xu, J. Q., and Arias, E. (2024). Mortality in the U nited S tates, 2023. Technical Report 521, National Center for Health Statistics, Hyattsville, MD. NCHS Data Brief, no 521. Hyattsville, MD: National Center for Health Statistics. 2024. DOI: ht...
2024 doi
-
[22]
Neal, R. M. (2000). M arkov chain sampling methods for D irichlet process mixture models. Journal of computational and graphical statistics , 9(2):249--265
2000
-
[23]
Pitman, J. (1996). Some developments of the B lackwell- M ac Q ueen urn scheme. Lecture Notes-Monograph Series , pages 245--267
1996
-
[24]
P., and Geller, M
Postman, M., Huchra, J. P., and Geller, M. J. (1986). Probes of large-scale structure in the corona borealis region. Astronomical Journal (ISSN 0004-6256), vol. 92, Dec. 1986, p. 1238-1247. , 92:1238--1247
1986
-
[25]
and Green, P
Richardson, S. and Green, P. J. (1997). On B ayesian analysis of mixtures with an unknown number of components (with discussion). Journal of the Royal Statistical Society Series B: Statistical Methodology , 59(4):731--792
1997
-
[26]
B., and Gelfand, A
Rodriguez, A., Dunson, D. B., and Gelfand, A. E. (2008). The nested D irichlet process. Journal of the American statistical Association , 103(483):1131--1154
2008
-
[27]
Roeder, K. (1990). Density estimation with confidence sets exemplified by superclusters and voids in the galaxies. Journal of the American Statistical Association , 85(411):617--624
1990
-
[28]
Samé, A., Ambroise, C., and Govaert, G. (2006). A classification em algorithm for binned data. Computational Statistics & Data Analysis , 51(2):466--480
2006
-
[29]
Schaeffer, K. (2024). U.s. centenarian population is projected to quadruple over the next 30 years. Pew Research Center . Accessed: YYYY-MM-DD
2024
-
[30]
A., da Silva, J
Shirazi, Z. A., da Silva, J. P. A. R., and de Souza, C. P. E. (2023). Parameter estimation for grouped data using EM and MCEM algorithms. Communications in Statistics - Simulation and Computation , 52(12):6513--6528
2023
-
[31]
H., Christensen, D., and Hjort, N
Simensen, O. H., Christensen, D., and Hjort, N. L. (2026). Random irregular histograms. Computational Statistics & Data Analysis , 220:108367
2026
-
[32]
Teh, Y., Jordan, M., Beal, M., and Blei, D. (2004). Sharing clusters among related groups: Hierarchical D irichlet processes. Advances in neural information processing systems , 17
2004
-
[33]
W., Jordan, M
Teh, Y. W., Jordan, M. I., Beal, M. J., and Blei, D. M. (2006). Hierarchical D irichlet processes. Journal of the american statistical association , 101(476):1566--1581
2006
-
[34]
Vallender, S. S. (1974). Calculation of the W asserstein distance between probability distributions on the line. Theory of Probability & Its Applications , 18(4):784--786
1974
-
[35]
Wasserman, L. (2006). All of nonparametric statistics . Springer Science & Business Media
2006
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.