A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending on the dataset.
Privacy Distillation: Reducing Re-identification Risk of Multimodal Diffusion Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Knowledge distillation in neural networks refers to compressing a large model or dataset into a smaller version of itself. We introduce Privacy Distillation, a framework that allows a text-to-image generative model to teach another model without exposing it to identifiable data. Here, we are interested in the privacy issue faced by a data provider who wishes to share their data via a multimodal generative model. A question that immediately arises is ``How can a data provider ensure that the generative model is not leaking identifiable information about a patient?''. Our solution consists of (1) training a first diffusion model on real data (2) generating a synthetic dataset using this model and filtering it to exclude images with a re-identifiability risk (3) training a second diffusion model on the filtered synthetic data only. We showcase that datasets sampled from models trained with privacy distillation can effectively reduce re-identification risk whilst maintaining downstream performance.
citation-role summary
citation-polarity summary
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending on the dataset.