Pith. sign in

REVIEW 3 cited by

Differentially Private Synthetic Data: Applied Evaluations and Enhancements

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.05537 v1 pith:Q3HBVQY2 submitted 2020-11-11 cs.LG cs.AIcs.CRcs.CY

classification cs.LGcs.AIcs.CRcs.CY
keywords dataprivatedifferentiallylearningmachineappliedmodelssynthetic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning practitioners frequently seek to leverage the most informative available data, without violating the data owner's privacy, when building predictive models. Differentially private data synthesis protects personal details from exposure, and allows for the training of differentially private machine learning models on privately generated datasets. But how can we effectively assess the efficacy of differentially private synthetic data? In this paper, we survey four differentially private generative adversarial networks for data synthesis. We evaluate each of them at scale on five standard tabular datasets, and in two applied industry scenarios. We benchmark with novel metrics from recent literature and other standard machine learning tools. Our results suggest some synthesizers are more applicable for different privacy budgets, and we further demonstrate complicating domain-based tradeoffs in selecting an approach. We offer experimental learning on applied machine learning scenarios with private internal data to researchers and practioners alike. In addition, we propose QUAIL, an ensemble-based modeling approach to generating synthetic data. We examine QUAIL's tradeoffs, and note circumstances in which it outperforms baseline differentially private supervised learning models under the same budget constraint.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are Data Experts Buying into Differentially Private Synthetic Data? Gathering Community Perspectives

    cs.HC 2024-12 conditional novelty 6.0 of 10

    Interviews with 17 data experts show skepticism toward differentially private synthetic data, a last-resort stance, and a demand for validation against real data.

  2. SafeSynthDP: Leveraging Large Language Models for Privacy-Preserving Synthetic Data Generation Using Differential Privacy

    cs.LG 2024-12 reject novelty 3.0 of 10

    A study of an LLM-based synthetic data pipeline with noise injection that claims differential privacy, evaluated on news classification, but with no valid privacy analysis.

  3. How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

    cs.CR 2025-12 conditional novelty 2.0 of 10

    A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.

Pith tools