Pith. sign in

REVIEW 2 major objections 1 minor 39 references

Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data

T0 review · 2 major / 1 minor · reviewed 2026-06-25 · grok-4.3

Pith's one-line read A Pitman-Yor posterior-predictive mechanism generates synthetic discrete data with instance-level differential privacy that strengthens as the discount parameter decreases.

desk verdict Pitman-Yor posterior predictive gives explicit regime-dependent DP guarantees and Wasserstein rates for discrete synthetic data, but the derivations need direct inspection. read the letter →

arxiv 2606.26073 v1 pith:VT4CRZ7B submitted 2026-06-24 math.ST stat.TH

classification math.STstat.TH
keywords syntheticdatagenerationdifferentialprivacyPitman-YorprocessBayesiannonparametricsdiscretedistributionsWassersteindistanceprivacy-utilitytradeoff
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper constructs a Bayesian nonparametric procedure for releasing synthetic samples from discrete confidential data. It places a Pitman-Yor process prior on the unknown discrete distribution, draws synthetic data from the posterior predictive, and derives differential privacy bounds that hold instance-wise. For positive discount parameter values the guarantee is (ε,δ)-differential privacy; for zero and negative values the guarantees strengthen when the released sample size satisfies explicit conditions. The same construction yields consistency of the synthetic empirical measure to the true data-generating distribution in 1-Wasserstein distance, together with explicit convergence rates that quantify how stronger privacy slows statistical utility.

What carries the argument

The Pitman-Yor posterior-predictive distribution, which produces synthetic data directly from the posterior after conditioning on the confidential sample and remains discrete almost surely.

What would settle it

A pair of neighboring confidential datasets and a choice of σ ∈ (0,1) for which the probability that the released synthetic sample falls in some set differs by more than the claimed (ε,δ) bound would falsify the privacy guarantee.

Watch

Extended reading notes

Core claim

The Pitman-Yor posterior-predictive mechanism provides an instance-level (ε,δ)-differential privacy guarantee for σ ∈ (0,1), stronger guarantees for σ = 0 and σ < 0 under suitable conditions on the released sample size, and proves consistency of the empirical distribution of the synthetic data in the 1-Wasserstein metric with explicit rates that make the privacy-utility tradeoff precise.

Load-bearing premise

The confidential data are modeled as a random sample from an unknown discrete distribution endowed with a Pitman-Yor process prior.

Editorial extensions

If this is right

  • For σ < 0 the mechanism reduces exactly to a parametric Dirichlet-Multinomial model and inherits stronger privacy.
  • Consistency in 1-Wasserstein distance holds for σ ≤ 0, with rates that degrade when privacy requirements force smaller released sample sizes.
  • The construction applies without prior knowledge of the number of categories because the Pitman-Yor random measure is discrete almost surely.
  • The explicit Wasserstein rates make the tension between privacy level and statistical accuracy quantifiable through the single choice of released sample size.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same posterior-predictive construction might be applied to other exchangeable nonparametric priors whose discount parameters control clustering behavior.
  • If the released sample size is chosen adaptively from the data, the privacy analysis would require additional arguments beyond those given for fixed sizes.
  • The Wasserstein consistency rates suggest that an optimal released size exists that balances the privacy constraint against convergence speed for any fixed privacy budget.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper develops a Bayesian nonparametric framework for privacy-preserving synthetic data generation for discrete data using the Pitman-Yor process prior. Synthetic data are drawn from the posterior-predictive distribution, and the authors establish differential privacy guarantees for different values of the discount parameter σ, with stronger guarantees for σ ≤ 0 under sample size restrictions. They also prove consistency of the empirical distribution of the synthetic data in the 1-Wasserstein metric and derive explicit convergence rates for σ ≤ 0.

Significance. If the claimed (ε,δ)-DP guarantees and explicit 1-Wasserstein consistency rates hold, the work provides a flexible nonparametric mechanism for discrete data with unknown or growing numbers of categories and makes the privacy-utility tradeoff precise via convergence rates.

major comments (2)
  1. [Abstract] Abstract: the instance-level (ε,δ)-DP claim for σ∈(0,1) and the stronger guarantees for σ=0, σ<0 (under released-sample-size conditions) rest on uninspectable derivations; without the proof steps it is impossible to verify correctness of the privacy analysis or the handling of the almost-surely discrete random measure across regimes.
  2. [Abstract] Abstract: the consistency and explicit convergence rates in 1-Wasserstein distance are stated only for σ≤0; the paper must confirm that the rates correctly quantify the tradeoff when the released sample size is restricted to obtain the stronger privacy guarantees.
minor comments (1)
  1. Clarify the precise definition of 'instance-level' differential privacy and how it relates to the standard neighboring-database definition used in the privacy analysis.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their careful reading of the manuscript and for highlighting the need for clearer verification of the privacy analysis and explicit confirmation of the privacy-utility tradeoff. We address each major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the instance-level (ε,δ)-DP claim for σ∈(0,1) and the stronger guarantees for σ=0, σ<0 (under released-sample-size conditions) rest on uninspectable derivations; without the proof steps it is impossible to verify correctness of the privacy analysis or the handling of the almost-surely discrete random measure across regimes.

    Authors: The complete derivations of the (ε,δ)-DP guarantees appear in Section 3. Theorem 3.1 establishes the instance-level guarantee for σ ∈ (0,1) by bounding the privacy loss of the Pitman-Yor posterior-predictive mechanism, explicitly accounting for the almost-sure discreteness of the random measure. Theorems 3.2 and 3.3 derive the stronger guarantees for σ = 0 and σ < 0 under the stated sample-size restrictions, again using the explicit form of the posterior-predictive distribution. If the intermediate steps remain difficult to follow, we will insert expanded proof sketches with additional intermediate inequalities in the revised manuscript. revision: partial

  2. Referee: [Abstract] Abstract: the consistency and explicit convergence rates in 1-Wasserstein distance are stated only for σ≤0; the paper must confirm that the rates correctly quantify the tradeoff when the released sample size is restricted to obtain the stronger privacy guarantees.

    Authors: Theorem 4.2 gives the explicit 1-Wasserstein rates for σ ≤ 0 as functions of the released sample size m. Section 4.3 then shows how the privacy constraints on m (required to obtain the stronger DP guarantees in Theorems 3.2–3.3) directly slow the convergence rate, thereby making the privacy-utility tradeoff precise. We will add a short clarifying sentence in the abstract and at the end of Section 4.3 to emphasize this dependence. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: derivations rely on external properties of Pitman-Yor process and standard DP definitions

full rationale

The paper's claims rest on establishing DP guarantees and 1-Wasserstein consistency rates via mathematical analysis of the Pitman-Yor posterior-predictive mechanism across σ regimes. No self-definitional reductions, fitted inputs renamed as predictions, or load-bearing self-citations appear in the abstract or described chain. Results invoke standard differential privacy definitions and known properties of the Pitman-Yor process (almost-sure discreteness, posterior-predictive sampling) without reducing the target quantities to quantities defined inside the paper by construction. The derivation is therefore self-contained against external benchmarks.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The framework rests on the standard assumption that discrete data arise from an unknown distribution equipped with a Pitman-Yor process prior; no new entities are postulated and the only free parameter is the discount σ whose regimes are analyzed separately.

free parameters (1)
  • discount parameter σ
    The model is analyzed in three fixed regimes of σ; the value is chosen by the analyst rather than fitted to the confidential data.
assumptions (2)
  • domain assumption Data are i.i.d. draws from an unknown discrete distribution equipped with a Pitman-Yor process prior
    Invoked in the first sentence of the abstract as the modeling choice that enables the posterior-predictive release mechanism.
  • standard math Standard definitions of (ε,δ)-differential privacy and 1-Wasserstein distance
    Used without re-derivation to state the privacy and utility results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data." pith.science (2026). https://pith.science/paper/VT4CRZ7B

@misc{pith2026260626073,
  author       = {Pith},
  title        = {Pith review of: Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VT4CRZ7B}},
  note         = {Machine review of arXiv:2606.26073}
}
abstract

Synthetic data generation is a powerful approach to privacy-preserving statistical analysis, where data-release mechanisms are governed by a privacy-utility tradeoff: they should provide privacy guarantees while preserving the statistical utility of confidential data. We develop a Bayesian nonparametric framework for private synthetic data generation tailored to discrete data. Specifically, the confidential data are modeled as a random sample from an unknown discrete distribution endowed with a Pitman-Yor process prior, and synthetic data are generated from the corresponding posterior-predictive distribution. Since the Pitman-Yor process defines an almost surely discrete random probability measure, the resulting mechanism is naturally suited to data with ties and settings involving a potentially large, unknown, or growing number of categories. We study differential privacy guarantees of the Pitman-Yor posterior-predictive mechanism across the three regimes of the discount parameter $\sigma\in(-\infty,1)$. For $\sigma\in(0,1)$, we establish an instance-level $(\varepsilon,\delta)$-differential privacy guarantee. For $\sigma=0$ and $\sigma<0$, corresponding respectively to the Dirichlet process prior and to a parametric Dirichlet-Multinomial model, stronger guarantees are obtained, under suitable conditions on the released sample size. We also investigate statistical utility, or informativity, of the released data via the expected $1$-Wasserstein distance between the empirical distribution of the synthetic data and the "true" data-generating distribution. For $\sigma<0$ and $\sigma=0$, we prove consistency of the empirical distribution in this metric and derive explicit convergence rates, making precise the privacy-utility tradeoff: stronger privacy guarantees impose more restrictive choices of the released sample size, slowing down convergence to the "true" data-generating distribution.

Figures

Figures reproduced from arXiv: 2606.26073 by the authors.

Figure 1
Figure 1. displays the results for θ ∈ {1, 10, 100}. In the three (ε, δ)-differentially private scenarios, the decay is close to the n −1/2 benchmark on the log-log scale, in agreement with Corollary 3. In the asymptotic ε-differential privacy regime, the convergence is slower, consis￾tently with Corollary 13. Although the asymptotic ε-DP curve can lie below some of the finite-δ curves for the sample sizes considered here, th… view at source ↗
Figure 3
Figure 3. Boxplots of the 1-Wasserstein dis￾tance between synthetic and confidential em￾pirical measures, over 100 independent runs. 5 Discussion We have proposed a Bayesian nonparametric framework for private synthetic data generation tailored to discrete data. The mechanism is model-aligned: synthetic observations are drawn directly from the posterior-predictive distribution of the Bayesian nonparametric model. Hence, under… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

39 extracted references · 1 canonical work pages

  1. [1]

    M., Ashmead, R., Cumings-Menon, R., Garfinkel, S., Heineck, M., Heiss, C., Johns, R., Kifer, D., Leclerc, P., Machanavajjhala, A., et al

    Abowd, J. M., Ashmead, R., Cumings-Menon, R., Garfinkel, S., Heineck, M., Heiss, C., Johns, R., Kifer, D., Leclerc, P., Machanavajjhala, A., et al. (2022). The 2020 census disclosure avoidance system topdown algorithm. arXiv , 2204.08986

  2. [2]

    and Gigli, N

    Ambrosio, L. and Gigli, N. (2013). A User's Guide to Optimal Transport , pages 1--155. Springer Berlin Heidelberg, Berlin, Heidelberg

  3. [3]

    Learning with privacy at scale

    Apple's Differential Privacy Team (2017). Learning with privacy at scale. Apple Machine Learning Journal , 1

  4. [4]

    Beraha, M., Favaro, S., and Rao, V. (2025). Mcmc for bayesian nonparametric mixture modeling under differential privacy. Journal of Computational and Graphical Statistics , 34(3):837--847

  5. [5]

    and Sheldon, D

    Bernstein, G. and Sheldon, D. R. (2019). Differentially private bayesian linear regression. In Advances in Neural Information Processing Systems , volume 32

  6. [6]

    Boedihardjo, M., Strohmer, T., and Vershynin, R. (2024). Private measures, random walks, and synthetic data. Probability theory and related fields , 189(1):569--611

  7. [7]

    Camerlenghi, F., Dolera, E., Favaro, S., and Mainini, E. (2026). Wasserstein posterior contraction rates in non-dominated bayesian nonparametric models. In Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , volume 62, pages 582--606. Institut Henri Poincar \'e

  8. [8]

    Dimitrakakis, C., Nelson, B., Zhang, Z., Mitrokotsa, A., and Rubinstein, B. I. P. (2017). Differential privacy for bayesian inference through posterior sampling. J. Mach. Learn. Res. , 18(1):343--381

Show all 39 references
  1. [9]

    Ding, B., Kulkarni, J., and Yekhanin, S. (2017). Collecting telemetry data privately. In Advances in Neural Information Processing Systems , pages 3571--3580

  2. [10]

    Dolera, E., Favaro, S., and Mainini, E. (2024). Strong posterior contraction rates via wasserstein dynamics. Probability Theory and Related Fields , 189(1):659--720

  3. [11]

    Donhauser, K., Abad, J., Hulkund, N., and Yang, F. (2024a). Privacy-preserving data release leveraging optimal transport and particle gradient descent. In Proceedings of the 41st International Conference on Machine Learning , ICML'24. JMLR.org

  4. [12]

    Donhauser, K., Lokna, J., Sanyal, A., Boedihardjo, M., H \"o nig, R., and Yang, F. (2024b). Certified private data release for sparse lipschitz functions. In International Conference on Artificial Intelligence and Statistics , pages 1396--1404. PMLR

  5. [13]

    and Haensch, A.-C

    Drechsler, J. and Haensch, A.-C. (2024). 30 years of synthetic data. Statistical Science , 39:221--242

  6. [14]

    Dwork, C., McSherry, F., Nissim, K., and Smith, A. (2006). Calibrating noise to sensitivity in private data analysis. In Theory of cryptography conference , pages 265--284. Springer

  7. [15]

    Erlingsson, \'U ., Pihur, V., and Korolova, A. (2014). Rappor: Randomized aggregatable privacy-preserving ordinal response. In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security

  8. [16]

    Ferguson, T. S. (1973). A Bayesian Analysis of Some Nonparametric Problems . The Annals of Statistics , 1(2):209 -- 230

  9. [17]

    and Guillin, A

    Fournier, N. and Guillin, A. (2015). On the rate of convergence in wasserstein distance of the empirical measure. Probability Theory and Related Fields , 162:707--738

  10. [18]

    and van der Vaart, A

    Ghosal, S. and van der Vaart, A. (2017). Fundamentals of Nonparametric Bayesian Inference . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press

  11. [19]

    Gotz, M., Machanavajjhala, A., Wang, G., Xiao, X., and Gehrke, J. (2011). Publishing search logs---a comparative study of privacy guarantees. IEEE transactions on knowledge and data engineering , 24(3):520--532

  12. [20]

    He, Y., Strohmer, T., Vershynin, R., and Zhu, Y. (2025). Differentially private low-dimensional synthetic data from high-dimensional datasets. Information and Inference: A Journal of the IMA , 14(1):iaae034

  13. [21]

    He, Y., Vershynin, R., and Zhu, Y. (2023). Algorithmically effective differentially private synthetic data. In The Thirty Sixth Annual Conference on Learning Theory , pages 3941--3968. PMLR

  14. [22]

    He, Y., Vershynin, R., and Zhu, Y. (2024). Online differentially private synthetic data generation. IEEE Transactions on Privacy , 1:19--30

  15. [23]

    R., and Savitsky, T

    Hu, J., Williams, M. R., and Savitsky, T. D. (2022). Mechanisms for global differential privacy under bayesian data synthesis. Statistica Sinica

  16. [24]

    E., Ghalebikesabi, S., and Holmes, C

    Jewson, J. E., Ghalebikesabi, S., and Holmes, C. C. (2023). Differentially private statistical inference through beta-divergence one posterior sampling. In Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S., editors, Advances in Neural Information Proces...

  17. [25]

    Machanavajjhala, A., Kifer, D., Abowd, J., Gehrke, J., and Vilhuber, L. (2008). Privacy: Theory meets practice on the map. In 2008 IEEE 24th international conference on data engineering , pages 277--286. IEEE

  18. [26]

    and Talwar, K

    McSherry, F. and Talwar, K. (2007). Mechanism design via differential privacy. In Annual IEEE Symposium on Foundations of Computer Science (FOCS) . IEEE

  19. [27]

    Perman, M., Pitman, J., and Yor, M. (1992). Size-biased sampling of P oisson point processes and excursions. Probability Theory and Related Fields , 92:21--39

  20. [28]

    Pitman, J. (1995). Exchangeable and partially exchangeable random partitions. Probability Theory and Related Fields , 102:145--158

  21. [29]

    Pitman, J. (1996). Some developments of the blackwell-macqueen urn scheme. Lecture Notes-Monograph Series , 30:245--267

  22. [30]

    Pitman, J. (2006). Combinatorial stochastic processes , volume 1875 of Lecture Notes in Mathematics . Springer-Verlag

  23. [31]

    and Yor, M

    Pitman, J. and Yor, M. (1997). The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator . The Annals of Probability , 25(2):855 -- 900

  24. [32]

    Roch, S. (2024). Modern Discrete Probability: An Essential Toolkit . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press

  25. [33]

    Rubin, D. B. (1993). Statistical disclosure limitation. Journal of official Statistics , 9(2):461--468

  26. [34]

    D., Williams, M

    Savitsky, T. D., Williams, M. R., and Hu, J. (2022). Bayesian pseudo posterior mechanism under asymptotic differential privacy. J. Mach. Learn. Res. , 23(1)

  27. [35]

    Soria-Comas, J., Domingo-Ferrer, J., Sanchez, D., and Megias, D. (2017). Individual differential privacy: A utility-preserving formulation of differential privacy guarantees. IEEE Transactions on Information Forensics and Security , 12(6):1418--1429

  28. [36]

    Census Bureau (2021)

    U.S. Census Bureau (2021). American community survey 1-year estimates public use microdata sample (pums). https://www.census.gov/programs-surveys/acs/data/experimental-data/2020-1-year-pums.html

  29. [37]

    Villani, C. (2008). Optimal Transport . Grundlehren der mathematischen Wissenschaften. Springer Berlin, Heidelberg, 1 edition. Hardcover ISBN: 978-3-540-71049-3, Softcover ISBN: 978-3-662-50180-1, eBook ISBN: 978-3-540-71050-9

  30. [38]

    and Zhou, S

    Wasserman, L. and Zhou, S. (2010). A statistical framework for differential privacy. Journal of the American Statistical Association , 105(489):375--389

  31. [39]

    and Bach, F

    Weed, J. and Bach, F. (2019). Sharp asymptotic and finite-sample rates of convergence of empirical measures in wasserstein distance. Bernoulli , 25(4A):2620--2648

Pith tools

Reviewed June 25, 2026 · model on record in the stance chip above.