REVIEW 2 major objections 1 minor 39 references
Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data
T0 review · 2 major / 1 minor · reviewed 2026-06-25 · grok-4.3
Pith's one-line read A Pitman-Yor posterior-predictive mechanism generates synthetic discrete data with instance-level differential privacy that strengthens as the discount parameter decreases.
desk verdict Pitman-Yor posterior predictive gives explicit regime-dependent DP guarantees and Wasserstein rates for discrete synthetic data, but the derivations need direct inspection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Pitman-Yor posterior-predictive distribution, which produces synthetic data directly from the posterior after conditioning on the confidential sample and remains discrete almost surely.
What would settle it
A pair of neighboring confidential datasets and a choice of σ ∈ (0,1) for which the probability that the released synthetic sample falls in some set differs by more than the claimed (ε,δ) bound would falsify the privacy guarantee.
Extended reading notes
Core claim
The Pitman-Yor posterior-predictive mechanism provides an instance-level (ε,δ)-differential privacy guarantee for σ ∈ (0,1), stronger guarantees for σ = 0 and σ < 0 under suitable conditions on the released sample size, and proves consistency of the empirical distribution of the synthetic data in the 1-Wasserstein metric with explicit rates that make the privacy-utility tradeoff precise.
Load-bearing premise
The confidential data are modeled as a random sample from an unknown discrete distribution endowed with a Pitman-Yor process prior.
Editorial extensions
If this is right
- For σ < 0 the mechanism reduces exactly to a parametric Dirichlet-Multinomial model and inherits stronger privacy.
- Consistency in 1-Wasserstein distance holds for σ ≤ 0, with rates that degrade when privacy requirements force smaller released sample sizes.
- The construction applies without prior knowledge of the number of categories because the Pitman-Yor random measure is discrete almost surely.
- The explicit Wasserstein rates make the tension between privacy level and statistical accuracy quantifiable through the single choice of released sample size.
Reading between the lines
- The same posterior-predictive construction might be applied to other exchangeable nonparametric priors whose discount parameters control clustering behavior.
- If the released sample size is chosen adaptively from the data, the privacy analysis would require additional arguments beyond those given for fixed sizes.
- The Wasserstein consistency rates suggest that an optimal released size exists that balances the privacy constraint against convergence speed for any fixed privacy budget.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Bayesian nonparametric framework for privacy-preserving synthetic data generation for discrete data using the Pitman-Yor process prior. Synthetic data are drawn from the posterior-predictive distribution, and the authors establish differential privacy guarantees for different values of the discount parameter σ, with stronger guarantees for σ ≤ 0 under sample size restrictions. They also prove consistency of the empirical distribution of the synthetic data in the 1-Wasserstein metric and derive explicit convergence rates for σ ≤ 0.
Significance. If the claimed (ε,δ)-DP guarantees and explicit 1-Wasserstein consistency rates hold, the work provides a flexible nonparametric mechanism for discrete data with unknown or growing numbers of categories and makes the privacy-utility tradeoff precise via convergence rates.
major comments (2)
- [Abstract] Abstract: the instance-level (ε,δ)-DP claim for σ∈(0,1) and the stronger guarantees for σ=0, σ<0 (under released-sample-size conditions) rest on uninspectable derivations; without the proof steps it is impossible to verify correctness of the privacy analysis or the handling of the almost-surely discrete random measure across regimes.
- [Abstract] Abstract: the consistency and explicit convergence rates in 1-Wasserstein distance are stated only for σ≤0; the paper must confirm that the rates correctly quantify the tradeoff when the released sample size is restricted to obtain the stronger privacy guarantees.
minor comments (1)
- Clarify the precise definition of 'instance-level' differential privacy and how it relates to the standard neighboring-database definition used in the privacy analysis.
Simulated Author's Rebuttal
We thank the referee for their careful reading of the manuscript and for highlighting the need for clearer verification of the privacy analysis and explicit confirmation of the privacy-utility tradeoff. We address each major comment below.
read point-by-point responses
-
Referee: [Abstract] Abstract: the instance-level (ε,δ)-DP claim for σ∈(0,1) and the stronger guarantees for σ=0, σ<0 (under released-sample-size conditions) rest on uninspectable derivations; without the proof steps it is impossible to verify correctness of the privacy analysis or the handling of the almost-surely discrete random measure across regimes.
Authors: The complete derivations of the (ε,δ)-DP guarantees appear in Section 3. Theorem 3.1 establishes the instance-level guarantee for σ ∈ (0,1) by bounding the privacy loss of the Pitman-Yor posterior-predictive mechanism, explicitly accounting for the almost-sure discreteness of the random measure. Theorems 3.2 and 3.3 derive the stronger guarantees for σ = 0 and σ < 0 under the stated sample-size restrictions, again using the explicit form of the posterior-predictive distribution. If the intermediate steps remain difficult to follow, we will insert expanded proof sketches with additional intermediate inequalities in the revised manuscript. revision: partial
-
Referee: [Abstract] Abstract: the consistency and explicit convergence rates in 1-Wasserstein distance are stated only for σ≤0; the paper must confirm that the rates correctly quantify the tradeoff when the released sample size is restricted to obtain the stronger privacy guarantees.
Authors: Theorem 4.2 gives the explicit 1-Wasserstein rates for σ ≤ 0 as functions of the released sample size m. Section 4.3 then shows how the privacy constraints on m (required to obtain the stronger DP guarantees in Theorems 3.2–3.3) directly slow the convergence rate, thereby making the privacy-utility tradeoff precise. We will add a short clarifying sentence in the abstract and at the end of Section 4.3 to emphasize this dependence. revision: yes
Circularity Check
No circularity: derivations rely on external properties of Pitman-Yor process and standard DP definitions
full rationale
The paper's claims rest on establishing DP guarantees and 1-Wasserstein consistency rates via mathematical analysis of the Pitman-Yor posterior-predictive mechanism across σ regimes. No self-definitional reductions, fitted inputs renamed as predictions, or load-bearing self-citations appear in the abstract or described chain. Results invoke standard differential privacy definitions and known properties of the Pitman-Yor process (almost-sure discreteness, posterior-predictive sampling) without reducing the target quantities to quantities defined inside the paper by construction. The derivation is therefore self-contained against external benchmarks.
Assumptions & free parameters
free parameters (1)
- discount parameter σ
assumptions (2)
- domain assumption Data are i.i.d. draws from an unknown discrete distribution equipped with a Pitman-Yor process prior
- standard math Standard definitions of (ε,δ)-differential privacy and 1-Wasserstein distance
Cite this review
Pith. "Pith review of Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data." pith.science (2026). https://pith.science/paper/VT4CRZ7B
@misc{pith2026260626073,
author = {Pith},
title = {Pith review of: Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/VT4CRZ7B}},
note = {Machine review of arXiv:2606.26073}
}
abstract
Synthetic data generation is a powerful approach to privacy-preserving statistical analysis, where data-release mechanisms are governed by a privacy-utility tradeoff: they should provide privacy guarantees while preserving the statistical utility of confidential data. We develop a Bayesian nonparametric framework for private synthetic data generation tailored to discrete data. Specifically, the confidential data are modeled as a random sample from an unknown discrete distribution endowed with a Pitman-Yor process prior, and synthetic data are generated from the corresponding posterior-predictive distribution. Since the Pitman-Yor process defines an almost surely discrete random probability measure, the resulting mechanism is naturally suited to data with ties and settings involving a potentially large, unknown, or growing number of categories. We study differential privacy guarantees of the Pitman-Yor posterior-predictive mechanism across the three regimes of the discount parameter $\sigma\in(-\infty,1)$. For $\sigma\in(0,1)$, we establish an instance-level $(\varepsilon,\delta)$-differential privacy guarantee. For $\sigma=0$ and $\sigma<0$, corresponding respectively to the Dirichlet process prior and to a parametric Dirichlet-Multinomial model, stronger guarantees are obtained, under suitable conditions on the released sample size. We also investigate statistical utility, or informativity, of the released data via the expected $1$-Wasserstein distance between the empirical distribution of the synthetic data and the "true" data-generating distribution. For $\sigma<0$ and $\sigma=0$, we prove consistency of the empirical distribution in this metric and derive explicit convergence rates, making precise the privacy-utility tradeoff: stronger privacy guarantees impose more restrictive choices of the released sample size, slowing down convergence to the "true" data-generating distribution.
Figures
Reference graph
Works this paper leans on
-
[1]
Abowd, J. M., Ashmead, R., Cumings-Menon, R., Garfinkel, S., Heineck, M., Heiss, C., Johns, R., Kifer, D., Leclerc, P., Machanavajjhala, A., et al. (2022). The 2020 census disclosure avoidance system topdown algorithm. arXiv , 2204.08986
-
[2]
and Gigli, N
Ambrosio, L. and Gigli, N. (2013). A User's Guide to Optimal Transport , pages 1--155. Springer Berlin Heidelberg, Berlin, Heidelberg
2013
-
[3]
Learning with privacy at scale
Apple's Differential Privacy Team (2017). Learning with privacy at scale. Apple Machine Learning Journal , 1
2017
-
[4]
Beraha, M., Favaro, S., and Rao, V. (2025). Mcmc for bayesian nonparametric mixture modeling under differential privacy. Journal of Computational and Graphical Statistics , 34(3):837--847
2025
-
[5]
and Sheldon, D
Bernstein, G. and Sheldon, D. R. (2019). Differentially private bayesian linear regression. In Advances in Neural Information Processing Systems , volume 32
2019
-
[6]
Boedihardjo, M., Strohmer, T., and Vershynin, R. (2024). Private measures, random walks, and synthetic data. Probability theory and related fields , 189(1):569--611
2024
-
[7]
Camerlenghi, F., Dolera, E., Favaro, S., and Mainini, E. (2026). Wasserstein posterior contraction rates in non-dominated bayesian nonparametric models. In Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , volume 62, pages 582--606. Institut Henri Poincar \'e
2026
-
[8]
Dimitrakakis, C., Nelson, B., Zhang, Z., Mitrokotsa, A., and Rubinstein, B. I. P. (2017). Differential privacy for bayesian inference through posterior sampling. J. Mach. Learn. Res. , 18(1):343--381
2017
Show all 39 references
-
[9]
Ding, B., Kulkarni, J., and Yekhanin, S. (2017). Collecting telemetry data privately. In Advances in Neural Information Processing Systems , pages 3571--3580
2017
-
[10]
Dolera, E., Favaro, S., and Mainini, E. (2024). Strong posterior contraction rates via wasserstein dynamics. Probability Theory and Related Fields , 189(1):659--720
2024
-
[11]
Donhauser, K., Abad, J., Hulkund, N., and Yang, F. (2024a). Privacy-preserving data release leveraging optimal transport and particle gradient descent. In Proceedings of the 41st International Conference on Machine Learning , ICML'24. JMLR.org
-
[12]
Donhauser, K., Lokna, J., Sanyal, A., Boedihardjo, M., H \"o nig, R., and Yang, F. (2024b). Certified private data release for sparse lipschitz functions. In International Conference on Artificial Intelligence and Statistics , pages 1396--1404. PMLR
-
[13]
and Haensch, A.-C
Drechsler, J. and Haensch, A.-C. (2024). 30 years of synthetic data. Statistical Science , 39:221--242
2024
-
[14]
Dwork, C., McSherry, F., Nissim, K., and Smith, A. (2006). Calibrating noise to sensitivity in private data analysis. In Theory of cryptography conference , pages 265--284. Springer
2006
-
[15]
Erlingsson, \'U ., Pihur, V., and Korolova, A. (2014). Rappor: Randomized aggregatable privacy-preserving ordinal response. In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security
2014
-
[16]
Ferguson, T. S. (1973). A Bayesian Analysis of Some Nonparametric Problems . The Annals of Statistics , 1(2):209 -- 230
1973
-
[17]
and Guillin, A
Fournier, N. and Guillin, A. (2015). On the rate of convergence in wasserstein distance of the empirical measure. Probability Theory and Related Fields , 162:707--738
2015
-
[18]
and van der Vaart, A
Ghosal, S. and van der Vaart, A. (2017). Fundamentals of Nonparametric Bayesian Inference . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press
2017
-
[19]
Gotz, M., Machanavajjhala, A., Wang, G., Xiao, X., and Gehrke, J. (2011). Publishing search logs---a comparative study of privacy guarantees. IEEE transactions on knowledge and data engineering , 24(3):520--532
2011
-
[20]
He, Y., Strohmer, T., Vershynin, R., and Zhu, Y. (2025). Differentially private low-dimensional synthetic data from high-dimensional datasets. Information and Inference: A Journal of the IMA , 14(1):iaae034
2025
-
[21]
He, Y., Vershynin, R., and Zhu, Y. (2023). Algorithmically effective differentially private synthetic data. In The Thirty Sixth Annual Conference on Learning Theory , pages 3941--3968. PMLR
2023
-
[22]
He, Y., Vershynin, R., and Zhu, Y. (2024). Online differentially private synthetic data generation. IEEE Transactions on Privacy , 1:19--30
2024
-
[23]
R., and Savitsky, T
Hu, J., Williams, M. R., and Savitsky, T. D. (2022). Mechanisms for global differential privacy under bayesian data synthesis. Statistica Sinica
2022
-
[24]
E., Ghalebikesabi, S., and Holmes, C
Jewson, J. E., Ghalebikesabi, S., and Holmes, C. C. (2023). Differentially private statistical inference through beta-divergence one posterior sampling. In Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S., editors, Advances in Neural Information Proces...
2023
-
[25]
Machanavajjhala, A., Kifer, D., Abowd, J., Gehrke, J., and Vilhuber, L. (2008). Privacy: Theory meets practice on the map. In 2008 IEEE 24th international conference on data engineering , pages 277--286. IEEE
2008
-
[26]
and Talwar, K
McSherry, F. and Talwar, K. (2007). Mechanism design via differential privacy. In Annual IEEE Symposium on Foundations of Computer Science (FOCS) . IEEE
2007
-
[27]
Perman, M., Pitman, J., and Yor, M. (1992). Size-biased sampling of P oisson point processes and excursions. Probability Theory and Related Fields , 92:21--39
1992
-
[28]
Pitman, J. (1995). Exchangeable and partially exchangeable random partitions. Probability Theory and Related Fields , 102:145--158
1995
-
[29]
Pitman, J. (1996). Some developments of the blackwell-macqueen urn scheme. Lecture Notes-Monograph Series , 30:245--267
1996
-
[30]
Pitman, J. (2006). Combinatorial stochastic processes , volume 1875 of Lecture Notes in Mathematics . Springer-Verlag
2006
-
[31]
and Yor, M
Pitman, J. and Yor, M. (1997). The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator . The Annals of Probability , 25(2):855 -- 900
1997
-
[32]
Roch, S. (2024). Modern Discrete Probability: An Essential Toolkit . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press
2024
-
[33]
Rubin, D. B. (1993). Statistical disclosure limitation. Journal of official Statistics , 9(2):461--468
1993
-
[34]
D., Williams, M
Savitsky, T. D., Williams, M. R., and Hu, J. (2022). Bayesian pseudo posterior mechanism under asymptotic differential privacy. J. Mach. Learn. Res. , 23(1)
2022
-
[35]
Soria-Comas, J., Domingo-Ferrer, J., Sanchez, D., and Megias, D. (2017). Individual differential privacy: A utility-preserving formulation of differential privacy guarantees. IEEE Transactions on Information Forensics and Security , 12(6):1418--1429
2017
-
[36]
Census Bureau (2021)
U.S. Census Bureau (2021). American community survey 1-year estimates public use microdata sample (pums). https://www.census.gov/programs-surveys/acs/data/experimental-data/2020-1-year-pums.html
2021
-
[37]
Villani, C. (2008). Optimal Transport . Grundlehren der mathematischen Wissenschaften. Springer Berlin, Heidelberg, 1 edition. Hardcover ISBN: 978-3-540-71049-3, Softcover ISBN: 978-3-662-50180-1, eBook ISBN: 978-3-540-71050-9
2008
-
[38]
and Zhou, S
Wasserman, L. and Zhou, S. (2010). A statistical framework for differential privacy. Journal of the American Statistical Association , 105(489):375--389
2010
-
[39]
and Bach, F
Weed, J. and Bach, F. (2019). Sharp asymptotic and finite-sample rates of convergence of empirical measures in wasserstein distance. Bernoulli , 25(4A):2620--2648
2019
Reviewed June 25, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.