Pith. sign in

REVIEW 4 cited by

Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.06470 v3 pith:RS36JDOG submitted 2022-12-13 cs.LG cs.CRstat.ML

classification cs.LGcs.CRstat.ML
keywords modelspubliclearningprivatedataprivacylargepretrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of differentially private machine learning can be boosted significantly by leveraging the transfer learning capabilities of non-private models pretrained on large public datasets. We critically review this approach. We primarily question whether the use of large Web-scraped datasets should be viewed as differential-privacy-preserving. We caution that publicizing these models pretrained on Web data as "private" could lead to harm and erode the public's trust in differential privacy as a meaningful definition of privacy. Beyond the privacy considerations of using public data, we further question the utility of this paradigm. We scrutinize whether existing machine learning benchmarks are appropriate for measuring the ability of pretrained models to generalize to sensitive domains, which may be poorly represented in public Web data. Finally, we notice that pretraining has been especially impactful for the largest available models -- models sufficiently large to prohibit end users running them on their own devices. Thus, deploying such models today could be a net loss for privacy, as it would require (private) data to be outsourced to a more compute-powerful third party. We conclude by discussing potential paths forward for the field of private learning, as public pretraining becomes more popular and powerful.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Membership Inference Attacks for Unseen Classes

    cs.LG 2025-06 conditional novelty 7.0 of 10

    In a new 'unseen class' setting for membership inference, quantile regression attacks outperform shadow model attacks, which fall to the level of a global-threshold baseline.

  2. Scaling Laws for Differentially Private Language Models

    cs.LG 2025-01 conditional novelty 7.0 of 10

    Differentially private language models obey scaling laws in which compute-optimal models are roughly 10-50x smaller than non-private Chinchilla-optimal models, with large batch sizes and rapid saturation of compute.

  3. Lower Bounds for Public-Private Learning under Distribution Shift

    cs.LG 2025-07 reject novelty 6.0 of 10

    For Gaussian mean estimation and linear regression with distribution shift, the paper claims that public data never provides complementary value: either public data alone suffices, or (for large shifts) private data a...

  4. How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

    cs.CR 2025-12 conditional novelty 2.0 of 10

    A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.

Pith tools