REVIEW 8 cited by
Winning the NIST Contest: A scalable and general approach to differentially private synthetic data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose a general approach for differentially private synthetic data generation, that consists of three steps: (1) select a collection of low-dimensional marginals, (2) measure those marginals with a noise addition mechanism, and (3) generate synthetic data that preserves the measured marginals well. Central to this approach is Private-PGM, a post-processing method that is used to estimate a high-dimensional data distribution from noisy measurements of its marginals. We present two mechanisms, NIST-MST and MST, that are instances of this general approach. NIST-MST was the winning mechanism in the 2018 NIST differential privacy synthetic data competition, and MST is a new mechanism that can work in more general settings, while still performing comparably to NIST-MST. We believe our general approach should be of broad interest, and can be adopted in future mechanisms for synthetic data generation.
Forward citations
Cited by 8 Pith papers
-
Workload-Preserving Differentially Private Synthetic Data for Causal Inference via Maximum-Entropy Calibration
Releasing DP synthetic data built from treatment-arm and outcome moments, plus noise-aware multiple imputation, gives calibrated ATE intervals at strict privacy budgets.
-
Differentially Private Explanations for Clusters
DPClustX privately selects the most informative attributes for each cluster and releases noisy histograms only for those attributes, providing differentially private explanations of clustering results.
-
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Post-processing fairness methods (ROC and EqOdds) provide the most stable fairness-utility trade-offs when training classifiers on DP synthetic tabular data, outperforming pre- and in-processing interventions.
-
Is API Access to LLMs Useful for Generating Private Synthetic Tabular Data?
API access to Gemini 1.0 Pro does not improve differentially private synthetic tabular data beyond established non-LLM baselines on the tested datasets and workloads.
-
Are Data Experts Buying into Differentially Private Synthetic Data? Gathering Community Perspectives
Interviews with 17 data experts show skepticism toward differentially private synthetic data, a last-resort stance, and a demand for validation against real data.
-
Quantitative Auditing of AI Fairness with Differentially Private Synthetic Data
Fairness metrics computed on differentially private synthetic data differ from real-data values by up to 0.32 for some measures, despite staying below 0.1 on average, across Adult, COMPAS, and Diabetes.
-
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.
-
Synthetic Tabular Data: Methods, Attacks and Defenses
A review of tabular synthetic data generation, privacy attacks, and defenses, whose central message is that synthetic data alone does not guarantee privacy.
Discussion (0). Continue with ORCID to comment.