Pith. sign in

REVIEW 1 cited by

Why Train More? Effective and Efficient Membership Inference via Memorization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08015 v1 pith:DHT7TKBX submitted 2023-10-12 cs.LG cs.CR

classification cs.LGcs.CR
keywords modelsmiasmemorizationsamplesshadowadversarydatadistribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Membership Inference Attacks (MIAs) aim to identify specific data samples within the private training dataset of machine learning models, leading to serious privacy violations and other sophisticated threats. Many practical black-box MIAs require query access to the data distribution (the same distribution where the private data is drawn) to train shadow models. By doing so, the adversary obtains models trained "with" or "without" samples drawn from the distribution, and analyzes the characteristics of the samples under consideration. The adversary is often required to train more than hundreds of shadow models to extract the signals needed for MIAs; this becomes the computational overhead of MIAs. In this paper, we propose that by strategically choosing the samples, MI adversaries can maximize their attack success while minimizing the number of shadow models. First, our motivational experiments suggest memorization as the key property explaining disparate sample vulnerability to MIAs. We formalize this through a theoretical bound that connects MI advantage with memorization. Second, we show sample complexity bounds that connect the number of shadow models needed for MIAs with memorization. Lastly, we confirm our theoretical arguments with comprehensive experiments; by utilizing samples with high memorization scores, the adversary can (a) significantly improve its efficacy regardless of the MIA used, and (b) reduce the number of shadow models by nearly two orders of magnitude compared to state-of-the-art approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Technical Report for the Forgotten-by-Design Project: Targeted Obfuscation for Machine Learning

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Per-sample gradient noise and exponential down-weighting of LIRA-identified vulnerable points reduce membership inference success on CIFAR-10 while keeping test accuracy near baseline.

Pith tools