Pith. sign in

REVIEW 2 cited by

Label-Only Membership Inference Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.14321 v3 pith:IWTXWHPR submitted 2020-07-28 cs.CR cs.LGstat.ML

classification cs.CRcs.LGstat.ML
keywords attacksmembershipinferencelabel-onlymodelconfidencedatamodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Membership inference attacks are one of the simplest forms of privacy leakage for machine learning models: given a data point and model, determine whether the point was used to train the model. Existing membership inference attacks exploit models' abnormal confidence when queried on their training data. These attacks do not apply if the adversary only gets access to models' predicted labels, without a confidence measure. In this paper, we introduce label-only membership inference attacks. Instead of relying on confidence scores, our attacks evaluate the robustness of a model's predicted labels under perturbations to obtain a fine-grained membership signal. These perturbations include common data augmentations or adversarial examples. We empirically show that our label-only membership inference attacks perform on par with prior attacks that required access to model confidences. We further demonstrate that label-only attacks break multiple defenses against membership inference attacks that (implicitly or explicitly) rely on a phenomenon we call confidence masking. These defenses modify a model's confidence scores in order to thwart attacks, but leave the model's predicted labels unchanged. Our label-only attacks demonstrate that confidence-masking is not a viable defense strategy against membership inference. Finally, we investigate worst-case label-only attacks, that infer membership for a small number of outlier data points. We show that label-only attacks also match confidence-based attacks in this setting. We find that training models with differential privacy and (strong) L2 regularization are the only known defense strategies that successfully prevents all attacks. This remains true even when the differential privacy budget is too high to offer meaningful provable guarantees.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Dynamics of Membership Privacy in Deep Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Per-sample membership vulnerability is established early in training, especially for hard-to-learn examples, and can be tracked on an FPR-TPR plane.

  2. Entangled Threats: A Unified Kill Chain Model for Quantum Machine Learning Security

    quant-ph 2025-07 conditional novelty 6.0 of 10

    The paper adapts kill chain methodology from classical IT security to quantum machine learning, organizing published QML attacks into a five-stage lifecycle with attacker roles, capabilities, and defenses.

Pith tools