Pith. sign in

REVIEW 2 major objections 1 minor 35 references

Bayesian Mixture Models for Histograms: with Applications to Large Datasets

T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read Bayesian mixture models recover underlying distributions from histogram bin counts alone via reversible jump MCMC.

desk verdict The paper puts reversible-jump MCMC on binned counts together with a Dirichlet-process layer for clustering multiple histograms; the identifiability of the mixtures from coarse bins is the main open question. read the letter →

arxiv 2606.24001 v1 pith:HLYQYXSI submitted 2026-06-22 stat.ME

classification stat.ME
keywords databayesianhistogramsmixtureapplicationsdistributionmethodmixtures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper develops a Bayesian method to estimate a population distribution when only aggregated histogram or frequency table data is available. It models the binned counts as coming from a mixture of normal distributions, places a prior on the number of components, and samples the posterior with reversible jump MCMC to handle both finite and infinite mixtures. The approach extends to multiple histograms by using a Dirichlet process to cluster them, sharing information across groups and yielding posterior probabilities of homogeneity. It reports strong performance on large-scale data, indicating that nonparametric Bayesian modeling can work directly with summarized inputs.

What carries the argument

Reversible jump MCMC applied to normal mixture models on histogram bin counts, with a Dirichlet process prior for simultaneous clustering of multiple histograms.

What would settle it

A simulation study in which the same dataset is available both in raw form and as histograms; the mixture recovered from the binned version should closely match the mixture fitted directly to the raw observations.

Watch

Extended reading notes

Core claim

The central claim is that placing a prior on the number of mixture components and performing reversible jump MCMC allows Bayesian inference of the underlying continuous distribution from binned data alone. The framework is extended to multiple histograms by using a Dirichlet process to cluster them, enabling information sharing across populations and providing a posterior probability to assess homogeneity between groups. Some theoretical results support the performance of the methodology.

Load-bearing premise

The data-generating process is well approximated by a mixture of normal distributions whose parameters can be recovered from the observed bin counts.

Editorial extensions

If this is right

  • The method recovers population distributions accurately from aggregated data without access to individual observations.
  • It scales to large datasets where processing raw records would be computationally prohibitive.
  • Multiple histograms can be clustered to share statistical strength and produce posterior probabilities of similarity between populations.
  • The modeling framework is stated to be flexible for extension to mixture families other than normals.
  • Theoretical support is provided for the consistency and performance of the binned-data inference.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • This approach could enable statistical analysis in privacy-restricted settings that release only histograms.
  • It suggests a route to density estimation when data arrives already summarized or streamed in aggregate form.
  • The clustering extension might apply to comparing distributions across sites or time periods without pooling raw records.
  • Direct comparison of binned versus raw inference on benchmark datasets would test how much information is lost in aggregation.
  • keywords:[
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript introduces a Bayesian nonparametric approach to inferring underlying distributions from histogram (binned) data by fitting finite or countably infinite mixtures of normals via reversible-jump MCMC, with an extension to simultaneous modeling and clustering of multiple histograms using a Dirichlet process; it claims strong empirical performance on large-scale data together with supporting theoretical results.

Significance. If the recovery of mixture parameters from binned counts proves stable, the method would offer a principled way to perform inference and clustering on aggregated or privacy-protected data, extending reversible-jump and Dirichlet-process techniques to histogram likelihoods.

major comments (2)
  1. [Theoretical results] Theoretical results section: the manuscript states that theoretical results support performance, yet provides no explicit argument or bound establishing identifiability or posterior consistency when bin widths are comparable to component scales or when components overlap; this directly underpins the claim that parameters can be recovered from binned counts alone.
  2. [Simulation studies] Simulation or application studies: the reported strong performance on large-scale data is not accompanied by recovery experiments that vary bin width relative to component variance or that quantify posterior multimodality under overlapping components; without such checks the central empirical claim remains unverified.
minor comments (1)
  1. [Abstract] The abstract does not specify the form of the prior on the number of components (e.g., Poisson or stick-breaking) or the precise reversible-jump proposal mechanism.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the thoughtful and constructive report. The comments identify important gaps in the theoretical justification and empirical validation of parameter recovery from binned data. We address each point below and will revise the manuscript to strengthen these aspects.

read point-by-point responses
  1. Referee: [Theoretical results] Theoretical results section: the manuscript states that theoretical results support performance, yet provides no explicit argument or bound establishing identifiability or posterior consistency when bin widths are comparable to component scales or when components overlap; this directly underpins the claim that parameters can be recovered from binned counts alone.

    Authors: The manuscript presents theoretical results on posterior consistency for the histogram likelihood under the reversible-jump MCMC scheme for normal mixtures, relying on standard conditions for Dirichlet process mixtures and the continuity of the binned likelihood. We acknowledge that these results do not include explicit identifiability bounds or consistency rates for the regime in which bin widths approach component scales or when components overlap substantially. We will revise the theoretical section to clarify the scope of the existing arguments and add a discussion of the additional conditions required in the overlapping or coarse-binning cases, together with references to related identifiability results for binned mixtures. revision: yes

  2. Referee: [Simulation studies] Simulation or application studies: the reported strong performance on large-scale data is not accompanied by recovery experiments that vary bin width relative to component variance or that quantify posterior multimodality under overlapping components; without such checks the central empirical claim remains unverified.

    Authors: The current simulation studies focus on large-scale data sets with fixed binning and demonstrate accurate recovery and clustering performance. We agree that systematic variation of bin width relative to component variance and explicit quantification of posterior multimodality under overlap would provide stronger verification of the central claim. We will add a new set of targeted simulation experiments addressing these regimes and report the corresponding recovery metrics and posterior diagnostics in the revised manuscript. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity detected; derivation is self-contained

full rationale

The manuscript introduces a Bayesian nonparametric mixture model for histogram data via reversible-jump MCMC, with a Dirichlet-process extension for multiple histograms. No equations, parameter-fitting steps, or self-citations are presented that reduce any claimed result to a definition or input by construction. The performance claims rest on the standard MCMC posterior rather than any self-referential prediction or renamed empirical pattern. The provided text contains no load-bearing self-citation chains or ansatz smuggling. This is the expected honest outcome for a methodological paper whose central procedure is externally verifiable against simulated or real binned data.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the assumption that normal mixtures are flexible enough for the target populations and that the reversible-jump sampler mixes adequately on binned likelihoods; no free parameters or invented entities are explicitly listed in the abstract.

assumptions (2)
  • domain assumption The underlying density belongs to the class of finite or countably infinite normal mixtures.
    Stated in the abstract as the modeling choice for the population distribution.
  • domain assumption Reversible jump MCMC can be implemented to sample from the posterior over both component number and parameters given only bin counts.
    The inference procedure is presented without further justification of mixing or convergence properties.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bayesian Mixture Models for Histograms: with Applications to Large Datasets." pith.science (2026). https://pith.science/paper/HLYQYXSI

@misc{pith2026260624001,
  author       = {Pith},
  title        = {Pith review of: Bayesian Mixture Models for Histograms: with Applications to Large Datasets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HLYQYXSI}},
  note         = {Machine review of arXiv:2606.24001}
}
read the original abstract

In many real-world scenarios, especially those involving privacy constraints or data summarization, data are available only in aggregated forms, such as histograms or frequency tables. This work introduces a novel Bayesian method for inferring the underlying population distribution by fitting a mixture model to binned data. While we focus on mixtures of normal distributions, the framework is flexible and can be extended to other distributional families. We place a prior distribution on the number of mixture components, accommodating both finite and countably infinite mixtures, and perform inference using reversible jump MCMC. The proposed approach demonstrates strong performance on large-scale data, showcasing the potential of nonparametric Bayesian modeling in practical applications. Furthermore, we extend the method to model multiple histograms simultaneously and cluster them using the Dirichlet process. This enables information sharing across populations and provides a principled posterior probability to assess homogeneity between groups. Some theoretical results supporting the performance of our proposed methodology are also discussed.

Figures

Figures reproduced from arXiv: 2606.24001 by the authors.

Figure 1
Figure 1. The results of a simulation study which varied the sample size and number of bins. The results are compared with the Wasserstein’s distance from the true data generating distribution. It is evident that a higher number of samples improves performance and, to a point, more bins also produces better performance. The same simulation results are shown in [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Similar to [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Timing comparisons with the DPM at four different sample sizes. As the number of samples goes up the DPM’s computation increases (dashed lines), however, our method is nearly invariant to the sample size, but its computation time increases as the number of bins increases (solid lines). Although the timing numbers look promising, the model accuracy is also of prime consideration [PITH_FULL_IMAGE:figures/full_fig_p01… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Comparisons of accuracy with the DPM model (dashed lines) at four different sample sizes. Clearly with sufficient data and number of bins our method (solid lines) achieves the accuracy of the DPM model. Using this model as the data generating distribution, the model in…
Figure 5
Figure 5. Figure 5: Timing comparisons with the DPM for the misspecified case using four different sample sizes. Clearly as the number of samples goes up the computation time (in minutes) for the DPM increases, however, our method is somewhat invariant to the sample size, but its computat…
Figure 6
Figure 6. Figure 6: Comparisons of accuracy with the DPM model for the misspecified case using four different sample sizes. Clearly with sufficient data and number of bins our method begins to match the accuracy of the DPM model. 4.2 Clustering Multiple Histograms We evaluate the performa…
Figure 7
Figure 7. Figure 7: The four distributions which generated the data for this simulation study. We specified 7 sample sizes and number of bins configurations, each defining a separate row in [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: A comparison of a barchart of the flight times less than 350, in black, compared with 100 posterior draws of the fitted mixture model pdf (utilizing 723 bins), in red. 0 50 100 150 200 250 300 350 0.0 0.2 0.4 0.6 0.8 1.0 Flight Time (in minutes) CDFs DP base measure 48…
Figure 9
Figure 9. Figure 9: Comparison of the DP posterior base measure to the mean posterior CDF using the model in (1) with 5 different bin sizes. In the left plot it appears that all the models are in fairly close agreement. The left plot is a zoomed-in view and shows some differences between …
Figure 10
Figure 10. Figure 10: COVID-19 mortality clustering by U.S. state and sex. There are clear spatial effects as well as differences in mortality by sex. No states have females and males in the same cluster. 0 20 40 60 80 100 Age at COVID−19 Mortality Cluster 1: mean mortality 70.7 Cluster 2:…
Figure 11
Figure 11. Figure 11: Posterior mean density for each cluster (fixed at the salso estimate), as shown in [PITH_FULL_IMAGE:figures/full_fig_p023_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 2 canonical work pages

  1. [1]

    Alston, C. L. and Mengersen, K. L. (2010). Allowing for the effect of data binning in a B ayesian normal mixture model. Computational Statistics and Data Analysis , 54(4):916--923

  2. [2]

    Argiento, R., Cremaschi, A., and Vannucci, M. (2020). Hierarchical normalized completely random measures to cluster grouped data. Journal of the American Statistical Association , 115(529):318--333

  3. [3]

    Beraha, M., Guglielmi, A., and Quintana, F. A. (2021). The semi-hierarchical D irichlet process and its application to clustering homogeneous distributions. Bayesian Analysis , 16(4):1187--1219

  4. [4]

    and Kim, J

    Billard, L. and Kim, J. (2017). Hierarchical clustering for histogram data. Wiley Interdisciplinary Reviews: Computational Statistics , 9(5):e1405

  5. [5]

    and Dunson, D

    Canale, A. and Dunson, D. B. (2011). B ayesian kernel mixtures for counts. Journal of the American Statistical Association , 106(496):1528--1539

  6. [6]

    Corradin, R., Canale, A., and Nipoti, B. (2021). BNPmix : An R package for B ayesian nonparametric modeling via P itman- Y or mixtures. Journal of Statistical Software , 100(15):1--33

  7. [7]

    B., Johnson, D

    Dahl, D. B., Johnson, D. J., and M \"u ller, P. (2022). Search algorithms and loss functions for B ayesian clustering. Journal of Computational and Graphical Statistics , 31(4):1189--1201

  8. [8]

    Denti, F., Camerlenghi, F., Guindani, M., and Mira, A. (2023). A common atoms model for the bayesian nonparametric analysis of nested data. Journal of the American Statistical Association , 118(541):405--416

Show all 35 references
  1. [9]

    Duan, Y., Guo, S., Wang, W., and M \"u ller, P. (2025). Immune profiling among colorectal cancer subtypes using dependent mixture models. Journal of the American Statistical Association , 120(550):671--684

  2. [10]

    and Denti, F

    D’Angelo, L. and Denti, F. (2026). A finite-infinite shared atoms nested model for the bayesian analysis of large grouped data sets. Bayesian Analysis , 21(1):105--138

  3. [11]

    Escobar, M. D. and West, M. (1995). B ayesian density estimation and inference using mixtures. Journal of the american statistical association , 90(430):577--588

  4. [12]

    Gau, S., de Dieu Tapsoba, J., and Lee, S. (2014). Bayesian approach for mixture models with grouped data. Computational Statistics , 29(5):1025--1043

  5. [13]

    and van der Vaart, A

    Ghosal, S. and van der Vaart, A. W. (2007). Convergence rates of posterior distributions for non-iid observations. Annals of Statistics , 35(1):192--223

  6. [14]

    T., Semenova, L., and Rudin, C

    Goh, S. T., Semenova, L., and Rudin, C. (2024). Sparse density trees and lists: An interpretable alternative to high-dimensional histograms. INFORMS Journal on Data Science

  7. [15]

    and Arabie, P

    Hubert, L. and Arabie, P. (1985). Comparing partitions. Journal of classification , 2(1):193--218

  8. [16]

    and James, L

    Ishwaran, H. and James, L. F. (2001). Gibbs sampling methods for stick-breaking priors. Journal of the American Statistical Association , 96(453):161--173

  9. [17]

    and Eilers, P

    Lambert, P. and Eilers, P. H. C. (2009). Bayesian density estimation from grouped continuous data. Computational Statistics & Data Analysis , 53(4):1388--1399

  10. [18]

    Lijoi, A., Pr \"u nster, I., and Rebaudo, G. (2023). Flexible clustering via hidden hierarchical dirichlet priors. Scandinavian Journal of Statistics , 50(1):213--234

  11. [19]

    Mart \' nez, A. F. and D \' az-Avalos, C. (2024). A model-based approach for clustering binned data. arXiv preprint arXiv:2409.07738

  12. [20]

    Miller, J. W. and Harrison, M. T. (2018). Mixture models with a prior on the number of components. Journal of the American Statistical Association , 113(521):340--356

  13. [21]

    L., Kochanek, K

    Murphy, S. L., Kochanek, K. D., Xu, J. Q., and Arias, E. (2024). Mortality in the U nited S tates, 2023. Technical Report 521, National Center for Health Statistics, Hyattsville, MD. NCHS Data Brief, no 521. Hyattsville, MD: National Center for Health Statistics. 2024. DOI: ht...

  14. [22]

    Neal, R. M. (2000). M arkov chain sampling methods for D irichlet process mixture models. Journal of computational and graphical statistics , 9(2):249--265

  15. [23]

    Pitman, J. (1996). Some developments of the B lackwell- M ac Q ueen urn scheme. Lecture Notes-Monograph Series , pages 245--267

  16. [24]

    P., and Geller, M

    Postman, M., Huchra, J. P., and Geller, M. J. (1986). Probes of large-scale structure in the corona borealis region. Astronomical Journal (ISSN 0004-6256), vol. 92, Dec. 1986, p. 1238-1247. , 92:1238--1247

  17. [25]

    and Green, P

    Richardson, S. and Green, P. J. (1997). On B ayesian analysis of mixtures with an unknown number of components (with discussion). Journal of the Royal Statistical Society Series B: Statistical Methodology , 59(4):731--792

  18. [26]

    B., and Gelfand, A

    Rodriguez, A., Dunson, D. B., and Gelfand, A. E. (2008). The nested D irichlet process. Journal of the American statistical Association , 103(483):1131--1154

  19. [27]

    Roeder, K. (1990). Density estimation with confidence sets exemplified by superclusters and voids in the galaxies. Journal of the American Statistical Association , 85(411):617--624

  20. [28]

    Samé, A., Ambroise, C., and Govaert, G. (2006). A classification em algorithm for binned data. Computational Statistics & Data Analysis , 51(2):466--480

  21. [29]

    Schaeffer, K. (2024). U.s. centenarian population is projected to quadruple over the next 30 years. Pew Research Center . Accessed: YYYY-MM-DD

  22. [30]

    A., da Silva, J

    Shirazi, Z. A., da Silva, J. P. A. R., and de Souza, C. P. E. (2023). Parameter estimation for grouped data using EM and MCEM algorithms. Communications in Statistics - Simulation and Computation , 52(12):6513--6528

  23. [31]

    H., Christensen, D., and Hjort, N

    Simensen, O. H., Christensen, D., and Hjort, N. L. (2026). Random irregular histograms. Computational Statistics & Data Analysis , 220:108367

  24. [32]

    Teh, Y., Jordan, M., Beal, M., and Blei, D. (2004). Sharing clusters among related groups: Hierarchical D irichlet processes. Advances in neural information processing systems , 17

  25. [33]

    W., Jordan, M

    Teh, Y. W., Jordan, M. I., Beal, M. J., and Blei, D. M. (2006). Hierarchical D irichlet processes. Journal of the american statistical association , 101(476):1566--1581

  26. [34]

    Vallender, S. S. (1974). Calculation of the W asserstein distance between probability distributions on the line. Theory of Probability & Its Applications , 18(4):784--786

  27. [35]

    Wasserman, L. (2006). All of nonparametric statistics . Springer Science & Business Media

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.